Robotics Papers

2024-10-31 · ICRA · 219 citations · club pick

EgoMimic: Scaling Imitation Learning via Egocentric Video

Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, Danfei Xu

Published at ICRA (the arXiv record still lists it as a preprint). 219 citations, 14 of them influential, as of the last refresh.

Abstract

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

End-to-end imitation learning has shown remarkable performance in learning complex manipulation tasks, but it remains brittle when facing new scenarios and tasks. Drawing on the recent success of Computer Vision and Natural Language Processing, we hypothesize that for learned policies to achieve broad generalization, we must dramatically scale up the training data size. While these adjacent domains benefit from Internet-sourced data, robotics lacks such an equivalent. <span id="page-0-0"></span> To scale up data…

SLIDE 2

What came before

Imitation Learning: Imitation Learning (IL) has been used to perform diverse and contact-rich manipulation tasks , , . Recent advancements in IL have led to the development of pixel-to-action IL models, which directly map raw visual inputs to low-level robot control , . These visual IL models have demonstrated impressive reactive policies ,

SLIDE 3

The method

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 4

What they measured

H1: EgoMimic is able to leverage human embodiment data to boost in-domain performance for complex manipulation tasks. H2: Human data helps EgoMimic generalize to new objects and scenes. H3: Given sufficient initial robot data, it is more valuable to collect additional human data than additional robot data. We select a set of long-horizon real world tasks to evaluate our

SLIDE 5

Where it breaks

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 6

One-line takeaway

Learns robot policies from egocentric human video via latent action inference + temporal alignment.

Assembled from the paper's own PDF, parsed with its layout intact, 62,595 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Presented at

Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “Learning Dexterous Manipulation from Egocentric Human Videos”. The arXiv title above is the record.
San Francisco, CA
Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
Listed on the event page as “X: https://x.com/TX\Leo\Wang/status/2059320921228546220 (Highly recommend)”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Learns robot policies from egocentric human video via latent action inference + temporal alignment.

Read next

Something wrong on this page? Open a correction.