2024-10-31 · ICRA · 219 citations · club pick
EgoMimic: Scaling Imitation Learning via Egocentric Video
Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, Danfei Xu
Published at ICRA (the arXiv record still lists it as a preprint). 219 citations, 14 of them influential, as of the last refresh.
Abstract
The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/
Ten-minute slide kit
Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.
SLIDE 1The problem
End-to-end imitation learning has shown remarkable performance in learning complex manipulation tasks, but it remains brittle when facing new scenarios and tasks. Drawing on the recent success of Computer Vision and Natural Language Processing, we hypothesize that for learned policies to achieve broad generalization, we must dramatically scale up the training data size. While these adjacent domains benefit from Internet-sourced data, robotics lacks such an equivalent. <span id="page-0-0"></span> To scale up data…
SLIDE 2What came before
Imitation Learning: Imitation Learning (IL) has been used to perform diverse and contact-rich manipulation tasks , , . Recent advancements in IL have led to the development of pixel-to-action IL models, which directly map raw visual inputs to low-level robot control , . These visual IL models have demonstrated impressive reactive policies ,
SLIDE 3The method
Not recoverable from the parsed text. Read this section in the paper yourself.
SLIDE 4What they measured
H1: EgoMimic is able to leverage human embodiment data to boost in-domain performance for complex manipulation tasks. H2: Human data helps EgoMimic generalize to new objects and scenes. H3: Given sufficient initial robot data, it is more valuable to collect additional human data than additional robot data. We select a set of long-horizon real world tasks to evaluate our
SLIDE 5Where it breaks
Not recoverable from the parsed text. Read this section in the paper yourself.
SLIDE 6One-line takeaway
Learns robot policies from egocentric human video via latent action inference + temporal alignment.
Assembled from the paper's own PDF, parsed with its layout intact, 62,595 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.
Presented at
Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “Learning Dexterous Manipulation from Egocentric Human Videos”. The arXiv title above is the record.
San Francisco, CA
Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
Listed on the event page as “X: https://x.com/TX\Leo\Wang/status/2059320921228546220 (Highly recommend)”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Learns robot policies from egocentric human video via latent action inference + temporal alignment.
Read next
2025-09-23
Ryan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa +5 · 30 citations
2025-09-04
Lawrence Y. Zhu, Pranav Kuppili, Ryan Punamiya +5 · 35 citations
2025-09-13
Yangcen Liu, Woo Chul Shin, Yunhai Han +3 · 27 citations
2026-04-08
Ryan Punamiya, Simar Kareer, Zeyi Liu +37 · 34 citations
2026-02-18
Ruijie Zheng, Dantong Niu, Yuqi Xie +12 · 71 citations