Robotics Papers

2026-02-18 · 71 citations · club pick

EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Castañeda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, Linxi Fan

No peer-reviewed venue on record yet. 71 citations, 5 of them influential, as of the last refresh.

Abstract

Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While prior work demonstrates human to robot transfer in constrained settings, it is unclear whether large scale human data can support fine grained, high degree of freedom dexterous manipulation. We present EgoScale, a human to dexterous manipulation transfer framework built on large scale egocentric human data. We train a Vision Language Action (VLA) model on over 20,854 hours of action labeled egocentric human video, more than 20 times larger than prior efforts, and uncover a log linear scaling law between human data scale and validation loss. This validation loss strongly correlates with downstream real robot performance, establishing large scale human data as a predictable supervision source. Beyond scale, we introduce a simple two stage transfer recipe: large scale human pretraining followed by lightweight aligned human robot mid training. This enables strong long horizon dexterous manipulation and one shot task adaptation with minimal robot supervision. Our final policy improves average success rate by 54% over a no pretraining baseline using a 22 DoF dexterous robotic hand, and transfers effectively to robots with lower DoF hands, indicating that large scale human motion provides a reusable, embodiment agnostic motor prior.

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Human behavior is one of the most scalable sources of data for learning physical intelligence. Humans routinely perform dexterous manipulation across diverse objects, environments, and task variations at a scale that far exceeds what can be collected through robot teleoperation. As robotic hardware continues to improve toward more human-like kinematics and dexterity, a natural question arises: *can human data serve as a primary training signal for dexterous robot manipulation?* Recent work shows that transfer from…

SLIDE 2

What came before

**Robot Learning from Human Data.** Human demonstrations have been widely used to scale robot learning, with early works leveraging human videos primarily for representation learning or intent inference \[17, 36, 15, 14, [45\]](#page-21-0). Subsequent approaches use human data to guide planning or high-level control while relying on robot demonstrations for low-level execution \[34, 43, 44, 18, 38, [39\]](#page-20-6). More recent methods exploit advances in egocentric sensing and 3D hand tracking to treat human…

SLIDE 3

The method

We aim to learn representations from large-scale egocentric human video that are directly useful for dexterous robot control. First, human demonstrations are noisy and lack paired robot actions. Second, human and robot embodiments differ substantially in kinematics and control interfaces. Our method (Figure 1\) addresses these challenges through two design

SLIDE 4

What they measured

We collect demonstrations on four bottles of different sizes, with 25 trajectories per bottle. <span id="page-6-0"></span> **(Task V)** *(Syringe) Syringe Liquid Transfer.* This is the most challenging task, requiring the robot to pick up a syringe, draw liquid from tube A, inject it into tube B, and discard the syringe into a trash can. The task involves long-horizon, multi-step reasoning, precise spatial alignment for fluid extraction and injection, and dexterous manipulation of the syringe plunger. **Evaluation…

SLIDE 5

Where it breaks

Importantly, the G1 is never trained from scratch. Instead, mid-training serves to align an already learned, human-derived manipulation representation with a new embodiment. As shown in Section 3, this approach yields substantially higher performance than training directly on G1 data alone, indicating that large-scale human pretraining provides a reusable and embodiment-agnostic motor prior that can be efficiently adapted to robots with different kinematics and hand

SLIDE 6

One-line takeaway

Constructs action-conditioned latent world models from human experience data for prediction, planning, and imagination-based control.

Assembled from the paper's own PDF, parsed with its layout intact, 68,812 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Figures worth putting on a slide

  • Figure 1: **EgoScale: Two-stage human-to-robot learning framework.** A flow-based Vision-Language-Action (VLA) policy is first pretrained on 20,854 hours of egocentric human videos
  • Figure 2: **Human Data Collection and Model Architecture.** (**Left**) Aligned human-robot mid-training data are collected using the same sensing setup as the robot. Vive trackers
  • Figure 3: **Post-Training Evaluation Tasks.** Five dexterous manipulation tasks used to evaluate post-training performance
  • Figure 4: **Main Experimental Results.** Comparison of Human Pre-train + Mid-Training, Human Pretraining, and No Pretraining across five dexterous manipulation tasks under two eval
  • Figure 5: **Scaling behavior of human pretraining.** *Left:* Human validation loss versus training steps for models pretrained with increasing amounts of egocentric human data (1k–
  • Figure 6: **Aligned mid-training enables emergent one-shot transfer.** During post-training, the policy is trained on only a single robot demonstration per task, together with alig

Presented at

Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “World Models from Human Experience”. The arXiv title above is the record.
San Francisco, CA
Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
San Francisco, CA
Why the club picked it. Constructs action-conditioned latent world models from human experience data for prediction, planning, and imagination-based control.

Read next

Something wrong on this page? Open a correction.