2025-07-16 · 131 citations · club pick
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, Hongxu Yin, Sifei Liu, Song Han, Yao Lu, Xiaolong Wang
No peer-reviewed venue on record yet. 131 citations, 3 of them influential, as of the last refresh.
Abstract
Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we explore training Vision-Language-Action (VLA) models using egocentric human videos. The benefit of using human videos is not only for their scale but more importantly for the richness of scenes and tasks. With a VLA trained on human video that predicts human wrist and hand actions, we can perform Inverse Kinematics and retargeting to convert the human actions to robot actions. We fine-tune the model using a few robot manipulation demonstrations to obtain the robot policy, namely EgoVLA. We propose a simulation benchmark called Ego Humanoid Manipulation Benchmark, where we design diverse bimanual manipulation tasks with demonstrations. We fine-tune and evaluate EgoVLA with Ego Humanoid Manipulation Benchmark and show significant improvements over baselines and ablate the importance of human data. Videos can be found on our website: https://rchalyang.github.io/EgoVLA
arXiv comment: More videos can be found on our website: https://rchalyang.github.io/EgoVLA
Ten-minute slide kit
Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.
Assembled from the paper's own PDF, parsed with its layout intact, 76,513 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.
Figures worth putting on a slide
- Figure 1: EgoVLA. Our vision-language-action model learns manipulation skills from egocentric human videos and transfers them to a bimanual humanoid robot. The top row illustrates
- Figure 2: EgoVLA takes visual history, language instruction, and action query token as input. The latent features are converted to human action with the action head. We use the wri
- Figure 3: Human Data Following insights from language model and vision-language model training, we emphasize the importance of dataset structure in driving model performance. We co
- Figure 4: Unified Action Space: MANO hand parameters are used as a shared action space for humans and robots. For robot hands, during training, optimized mano parameters produce th
- Figure 5: Task Visualization. All simulated tasks with predicted wrist trajs from EgoVLA.
- Figure 6: Visual Instruction Following. Top: original HOI4D samples with corresponding language instructions. Red lines indicate ground-truth human wrist trajectories, and green li
Presented at
Read next
Something wrong on this page? Open a correction.