Robotics Papers

2026-05-24 · 9 citations · club pick

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos

No peer-reviewed venue on record yet. 9 citations, 1 of them influential, as of the last refresh.

Abstract

Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challenging due to the embodiment gap between human and robot in both visual appearance and kinematics. We present HumanEgo, a framework that bridges the embodiment gap by lifting each human demonstration to an entity-level representation of hand-object interaction, and training a flow matching policy with dense auxiliary objectives that amplify supervision from every trajectory. HumanEgo is robot-data-free, hardware-agnostic, data-efficient, and zero-shot human-to-robot transferable. With only 30 minutes of human videos per task, HumanEgo achieves 92.5% average success across four real-world tasks (75% with just 15 minutes), outperforms matched-time robot teleoperation by 41%, and robustly transfers zero-shot across novel robots, cameras, and environments. We release HumanEgo as an easy-to-use, open-source framework for learning robot policies directly from human data: https://github.com/TX-Leo/HumanEgo

arXiv comment: Project page: https://humanego-ai.github.io

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

State-of-the-art manipulation policies require hundreds to thousands of task-specific robot demonstrations \[1, 2, 3, 4, 5, [6\]](#page-8-0), which are costly, time-consuming, and inconvenient to collect. Human egocentric video offers a much cheaper and more accessible alternative: with a head-mounted camera [\[7\]](#page-8-0), a single person can collect task demonstrations anywhere, in minutes. We pursue a more direct goal: learning deployable manipulation policies from only minutes of human egocentric…

SLIDE 2

What came before

Recent years have produced rich large-scale egocentric and hand–object interaction datasets \[34, 35, 15, 36, 37, 38, 39, [14\]](#page-9-0) that provide the data foundation for learning manipulation from human video. Building on this foundation, one line of work scales up generalist policies and world models \[12, 13, 40, 41, [42\]](#page-11-0) that learn embodiment-agnostic representations from massive corpora, yet deployment demands enormous compute and per-task robot post-training. Another line cotrains \[8, 9,…

SLIDE 3

The method

Fig. 14: Hand tracking comparison on Serve Bread (45 demonstrations, ∼45 k frames). *Top— Smoothness:* per-frame jerk of the gripper midpoint (translational and angular) and of all 21 keypoints (lower is better, log scale). *Bottom—Accuracy vs. Aria-MPS:* per-keypoint shape error after Procrustes alignment, residual rotation error after subtracting the systematic frame offset, and fraction of frames with a valid hand detection. **Monocular RGB Stereo + IMU** Fig. 15: Hand Tracking Method Study. ICT consumes 3D…

SLIDE 4

What they measured

**HumanEgo-30 HumanEgo-15 ACT (Robot Teleop) SPOT ZeroMimic Track2Act PointPolicy EgoZero** Fig. 4: Overall Real-World Evaluation. Real-world success rate (%) for each method across all four tasks. HumanEgo with 30 min of data achieves the highest success rate on every task, demonstrating consistent improvements over both human-video baselines and robot teleoperation methods. HumanEgo achieves the highest success rate on every single

SLIDE 5

Where it breaks

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 6

One-line takeaway

Code: https://github.com/TX-Leo/HumanEgo

Assembled from the paper's own PDF, parsed with its layout intact, 119,670 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Presented at

Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
San Francisco, CA
Why the club picked it. Code: https://github.com/TX-Leo/HumanEgo

Read next

Something wrong on this page? Open a correction.