2026-02-18 · 71 citations · club pick
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Castañeda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, Linxi Fan
No peer-reviewed venue on record yet. 71 citations, 5 of them influential, as of the last refresh.
Abstract
Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While prior work demonstrates human to robot transfer in constrained settings, it is unclear whether large scale human data can support fine grained, high degree of freedom dexterous manipulation. We present EgoScale, a human to dexterous manipulation transfer framework built on large scale egocentric human data. We train a Vision Language Action (VLA) model on over 20,854 hours of action labeled egocentric human video, more than 20 times larger than prior efforts, and uncover a log linear scaling law between human data scale and validation loss. This validation loss strongly correlates with downstream real robot performance, establishing large scale human data as a predictable supervision source. Beyond scale, we introduce a simple two stage transfer recipe: large scale human pretraining followed by lightweight aligned human robot mid training. This enables strong long horizon dexterous manipulation and one shot task adaptation with minimal robot supervision. Our final policy improves average success rate by 54% over a no pretraining baseline using a 22 DoF dexterous robotic hand, and transfers effectively to robots with lower DoF hands, indicating that large scale human motion provides a reusable, embodiment agnostic motor prior.
Ten-minute slide kit
Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.
SLIDE 1The problem
Human behavior is one of the most scalable sources of data for learning physical intelligence. Humans routinely perform dexterous manipulation across diverse objects, environments, and task variations at a scale that far exceeds what can be collected through robot teleoperation. As robotic hardware continues to improve toward more human-like kinematics and dexterity, a natural question arises: *can human data serve as a primary training signal for dexterous robot manipulation?* Recent work shows that transfer from…
SLIDE 2What came before
**Robot Learning from Human Data.** Human demonstrations have been widely used to scale robot learning, with early works leveraging human videos primarily for representation learning or intent inference \[17, 36, 15, 14, [45\]](#page-21-0). Subsequent approaches use human data to guide planning or high-level control while relying on robot demonstrations for low-level execution \[34, 43, 44, 18, 38, [39\]](#page-20-6). More recent methods exploit advances in egocentric sensing and 3D hand tracking to treat human…
SLIDE 3The method
We aim to learn representations from large-scale egocentric human video that are directly useful for dexterous robot control. First, human demonstrations are noisy and lack paired robot actions. Second, human and robot embodiments differ substantially in kinematics and control interfaces. Our method (Figure 1\) addresses these challenges through two design
SLIDE 4What they measured
We collect demonstrations on four bottles of different sizes, with 25 trajectories per bottle. <span id="page-6-0"></span> **(Task V)** *(Syringe) Syringe Liquid Transfer.* This is the most challenging task, requiring the robot to pick up a syringe, draw liquid from tube A, inject it into tube B, and discard the syringe into a trash can. The task involves long-horizon, multi-step reasoning, precise spatial alignment for fluid extraction and injection, and dexterous manipulation of the syringe plunger. **Evaluation…
SLIDE 5Where it breaks
Importantly, the G1 is never trained from scratch. Instead, mid-training serves to align an already learned, human-derived manipulation representation with a new embodiment. As shown in Section 3, this approach yields substantially higher performance than training directly on G1 data alone, indicating that large-scale human pretraining provides a reusable and embodiment-agnostic motor prior that can be efficiently adapted to robots with different kinematics and hand
SLIDE 6One-line takeaway
Constructs action-conditioned latent world models from human experience data for prediction, planning, and imagination-based control.
Assembled from the paper's own PDF, parsed with its layout intact, 68,812 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.
Figures worth putting on a slide
- Figure 1: **EgoScale: Two-stage human-to-robot learning framework.** A flow-based Vision-Language-Action (VLA) policy is first pretrained on 20,854 hours of egocentric human videos
- Figure 2: **Human Data Collection and Model Architecture.** (**Left**) Aligned human-robot mid-training data are collected using the same sensing setup as the robot. Vive trackers
- Figure 3: **Post-Training Evaluation Tasks.** Five dexterous manipulation tasks used to evaluate post-training performance
- Figure 4: **Main Experimental Results.** Comparison of Human Pre-train + Mid-Training, Human Pretraining, and No Pretraining across five dexterous manipulation tasks under two eval
- Figure 5: **Scaling behavior of human pretraining.** *Left:* Human validation loss versus training steps for models pretrained with increasing amounts of egocentric human data (1k–
- Figure 6: **Aligned mid-training enables emergent one-shot transfer.** During post-training, the policy is trained on only a single robot demonstration per task, together with alig
Presented at
Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “World Models from Human Experience”. The arXiv title above is the record.
San Francisco, CA
Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
San Francisco, CA
Why the club picked it. Constructs action-conditioned latent world models from human experience data for prediction, planning, and imagination-based control.
Read next
2025-12-27
Simar Kareer, Karl Pertsch, James Darpinian +5 · 39 citations
2026-06-15
Yuvan Sharma, Dantong Niu, Anirudh Pai +16
2026-06-10
Yangcen Liu, Shuo Cheng, Xinchen Yin +6 · 5 citations
2026-02-06
Shenyuan Gao, William Liang, Kaiyuan Zheng +27 · 96 citations
2024-10-31
Simar Kareer, Dhruv Patel, Ryan Punamiya +5 · 219 citations