Robotics Papers

2025-12-27 · 39 citations · club pick

Emergence of Human to Robot Transfer in Vision-Language-Action Models

Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, Suraj Nair

No peer-reviewed venue on record yet. 39 citations, 4 of them influential, as of the last refresh.

Abstract

Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can come from human videos, which cover diverse real-world situations and are easy to obtain. However, it is difficult to train VLAs with human videos alone, and establishing a mapping between humans and robots requires manual engineering and presents a major research challenge. Drawing inspiration from advances in large language models, where the ability to learn from diverse supervision emerges with scale, we ask whether a similar phenomenon holds for VLAs that incorporate human video data. We introduce a simple co-training recipe, and find that human-to-robot transfer emerges once the VLA is pre-trained on sufficient scenes, tasks, and embodiments. Our analysis suggests that this emergent capability arises because diverse pretraining produces embodiment-agnostic representations for human and robot data. We validate these findings through a series of experiments probing human to robot skill transfer and find that with sufficiently diverse robot pre-training our method can nearly double the performance on generalization settings seen only in human data.

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Human knowledge provides the foundation to instill physical intelligence in robots. This manifests in many forms, from bootstrapping robot policies with human generated text and images via vision-language models, to mimicking human generated actions via robot teleoperation. While such techniques *indirectly* imbue the model with human experience, the right recipe to directly learn from human experience, for instance by watching a video of someone perform a task, remains an active area of research \[9, 2, 31, 5,…

SLIDE 2

What came before

Learning manipulation from human video has received significant attention due to its potential scalability. Over the years, advances have been made to leverage this data more directly for policy learning. Early works in this field leveraged human video data to train stronger vision encoders, which can improve downstream policy learning \[31, 30,

SLIDE 3

The method

The x-axis represents the diversity of the pre-training robot dataset, and the yellow and blue lines show the finetuning performance with and without human embodiment data. While both increase, the gain from leveraging human data only appears beyond a certain pre-training scale. We evaluate on a suite of four generalization scenarios shown only in the human data. *Abstract*—Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to…

SLIDE 4

What they measured

<span id="page-4-0"></span>To test whether π0.<sup>5</sup> + ego can generalize to new concepts from egocentric human data, we construct a suite of "generalization" scenarios that have limited coverage in robot data, but present in human data. These scenarios span generalizing to new scenes, objects and tasks. We begin our study by understanding whether our recipe can enable transfer to these new settings. Then, we validate our core hypothesis, which is that this transfer is an emergent property of diverse VLA…

SLIDE 5

Where it breaks

We study the emergence of human to robot transfer in our proposed recipe π0.<sup>5</sup> +ego. We find that with limited pretraining diversity, VLAs fail to transfer knowledge from human data, but as pretraining diversity grows past a critical threshold, transfer emerges. While our recipes leverage vast datasets of robot teleoperation data in pretraining, we ultimately only use 10s of hours of human data, and this data is collected in an episodic

SLIDE 6

One-line takeaway

Jointly optimizes robot morphology, interfaces, and teleoperation pipelines to improve scalability and reduce human demonstration cost.

Assembled from the paper's own PDF, parsed with its layout intact, 64,152 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Presented at

Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “Human-Robot Co-Design for Scalable Data Collection”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Jointly optimizes robot morphology, interfaces, and teleoperation pipelines to improve scalability and reduce human demonstration cost.

Read next

2024-10-31
Kevin Black, Noah Brown, Danny Driess +21 · 2,461 citations
2024-06-13
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti +15 · 3,113 citations
2025-04-22
Physical Intelligence, Kevin Black, Noah Brown +33 · 1,625 citations

Something wrong on this page? Open a correction.