Robotics Papers

2026-04-08 · 34 citations · club pick

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

Ryan Punamiya, Simar Kareer, Zeyi Liu, Josh Citron, Ri-Zhao Qiu, Xiongyi Cai, Alexey Gavryushin, Jiaqi Chen, Davide Liconti, Lawrence Y. Zhu, Patcharapong Aphiwetsa, Baoyu Li, Aniketh Cheluva, Pranav Kuppili, Yangcen Liu, Dhruv Patel, Aidan Gao, Hye-Young Chung, Ryan Co, Renee Zbizika, Jeff Liu, Xiaomeng Xu, Haoyu Xiong, Geng Chen, Sebastiano Oliani, Wenkai Xuan, Chenyu Yang, Xi Wang, James Fort, Richard Newcombe, Josh Gao, Jason Chong, Garrett Matsuda, Aseem Doriwala, Marc Pollefeys, Robert Katzschmann, Xiaolong Wang, Shuran Song, Judy Hoffman, Danfei Xu

No peer-reviewed venue on record yet. 34 citations, 3 of them influential, as of the last refresh.

Abstract

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limited in scope, difficult to extend, and fragmented across institutions. We introduce EgoVerse, a collaborative platform for human data-driven robot learning that unifies data collection, processing, and access under a shared framework, enabling contributions from individual researchers, academic labs, and industry partners. The current release includes 1,362 hours (80k episodes) of human demonstrations spanning 1,965 tasks, 240 scenes, and 2,087 unique demonstrators, with standardized formats, manipulation-relevant annotations, and tooling for downstream learning. Beyond the dataset, we conduct a large-scale study of human-to-robot transfer with experiments replicated across multiple labs, tasks, and robot embodiments under shared protocols. We find that policy performance generally improves with increased human data, but that effective scaling depends on alignment between human data and robot learning objectives. Together, the dataset, platform, and study establish a foundation for reproducible progress in human data-driven robot learning. Videos and additional information can be found at https://egoverse.ai/

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Recent progress in robot learning has shown that scaling data is a powerful driver of generalization \[5, 26, 31, [39\]](#page-10-2). Large-scale imitation learning has enabled policies to handle broader task distributions, more visual variation, and longer horizons, echoing trends seen in large vision and language models. However, unlike those domains, robot learning faces a fundamental bottleneck – collecting robot demonstrations requires physical hardware, expert teleoperation, and controlled

SLIDE 2

What came before

<span id="page-2-0"></span> Datasets of Human Activities. Large-scale human activity datasets such as Something-Something V2 [\[19\]](#page-9-4), Ego4D [\[20\]](#page-9-5), HOI4D [\[36\]](#page-10-6), EgoExo4D [\[21\]](#page-9-6), and Epic-Kitchens [\[12\]](#page-9-7) capture rich human behavior across diverse environments. However, they are not designed for robot

SLIDE 3

The method

To enable joint training across diverse embodiments, we adopt an encoder–decoder architecture with shallow, modality-specific stems [\[49\]](#page-11-10). Image observations are processed by a ResNet-18 [\[23\]](#page-9-15) backbone, while proprioceptive inputs are encoded with an MLP, before being tokenized into a shared space via learned query attention. A shared vision stem processes egocentric RGB observations from both human and robot embodiments, while separate stems handle robotspecific wrist cameras and…

SLIDE 4

What they measured

Evaluation is performed on four representative Flagship tasks shown in Fig. 8. We evaluate both in-domain (ID) settings, where task layouts match robot training data, and out-of-domain (OOD) settings with unseen objects and environments. For each method, we perform 20 ID and 20 OOD rollouts per task, with randomized initial conditions. Performance is measured using task-specific subtask metrics, including grasps, placements, intermediate manipulations, and full task

SLIDE 5

Where it breaks

Our study mainly focused on human-and-robot co-training. Future work should conduct broader algorithmic exploration (e.g., pre-train and fine-tune). Moreover, the scene and demonstrator diversity experiments rely exclusively on offline

SLIDE 6

One-line takeaway

Pretrains behavior foundation models on massive internet-scale human motion datasets for downstream robotic policy transfer.

Assembled from the paper's own PDF, parsed with its layout intact, 116,531 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Figures worth putting on a slide

  • Figure 1: A color microscopy image showing the distribution of EgoVerse-I and EgoVerse-II in a colored collagen structure. The collagen is colored in the top right corner with red

Presented at

Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “Internet-Scale Human Motion Pretraining for Robotics”. The arXiv title above is the record.
San Francisco, CA
Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
Listed on the event page as “X: https://x.com/TX\Leo\Wang/status/2059320921228546220 (Highly recommend)”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Pretrains behavior foundation models on massive internet-scale human motion datasets for downstream robotic policy transfer.

Read next

2021-10-13
Kristen Grauman, Andrew Westbury, Eugene Byrne +82 · 1,994 citations
2023-08-24
Jakob Engel, Kiran Somasundaram, Michael Goesele +71 · 272 citations
2025-09-04
Lawrence Y. Zhu, Pranav Kuppili, Ryan Punamiya +5 · 35 citations
2025-03-17
Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng +12 · 95 citations

Something wrong on this page? Open a correction.