Robotics Papers

2026-02-06 · 96 citations · club pick

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Shenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik, Seonghyeon Ye, Sihyun Yu, Wei-Cheng Tseng, Yuzhu Dong, Kaichun Mo, Chen-Hsuan Lin, Qianli Ma, Seungjun Nah, Loic Magne, Jiannan Xiang, Yuqi Xie, Ruijie Zheng, Dantong Niu, You Liang Tan, K. R. Zentner, George Kurian, Suneel Indupuru, Pooya Jannaty, Jinwei Gu, Jun Zhang, Jitendra Malik, Pieter Abbeel, Ming-Yu Liu, Yuke Zhu, Joel Jang, Linxi "Jim" Fan

No peer-reviewed venue on record yet. 96 citations, 11 of them influential, as of the last refresh.

Abstract

Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. Our data mixture represents the largest video dataset to date for world model pretraining, spanning a wide range of daily scenarios with diverse objects and skills. To address the scarcity of action labels, we introduce continuous latent actions as unified proxy actions, enhancing interaction knowledge transfer from unlabeled videos. After post-training on small-scale target robot data, DreamDojo demonstrates a strong understanding of physics and precise action controllability. We also devise a distillation pipeline that accelerates DreamDojo to a real-time speed of 10.81 FPS and further improves context consistency. Our work enables several important applications based on generative world models, including live teleoperation, policy evaluation, and model-based planning. Systematic evaluation on multiple challenging out-of-distribution (OOD) benchmarks verifies the significance of our method for simulating open-world, contact-rich tasks, paving the way for general-purpose robot world models.

arXiv comment: Project page: https://dreamdojo-world.github.io/

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

World models, which predict futures based on actions, have emerged as a key component in the development of generalist robots \(Hu et al., 2023; LeCun, 2022; Richens et al., 2025; Sutton, 1991\). Recent advances in video generation \(Ali et al., 2025; Wan et al., 2025\) have driven video world models, in which future states are represented as video frames \(Ball et al., 2025; Russell et al., 2025; Sun et al., 2025\). However, they primarily plateau at discrete controls, while the high-dimensional action spaces for…

SLIDE 2

What came before

**World model.** World models can simulate world transitions in response to actions, which have been proven critical for developing intelligent agents \(Alonso et al., 2024; Ha and Schmidhuber, 2018; Hafner et al., 2025; Richens et al., 2025\). However, existing models are typically trained and evaluated in in-distribution settings, leaving it unclear whether these models can truly facilitate planning in unseen scenarios. Another thread of research focuses on world model pretraining from internet-scale videos to…

SLIDE 3

The method

Different from interactive games with discrete inputs \(Parker-Holder et al., 2024\), achieving genuine controllability for robot actions presents more challenges due to its high dimensionality and contact-rich nature. To realize precise action following, we propose two improvements based on the original architecture. First, instead of using the absolute robot joint poses, we transform them into relative actions by rebaselining the inputs with the pose at the beginning of each latent frame (*i.e*., every 4…

SLIDE 4

What they measured

In this section, we conduct extensive experiments to demonstrate DreamDojo's strengths. The dimension of the latent action is 32. The model has 24 encoder blocks for latent action extraction and 24 decoder blocks for forward dynamics prediction. It is trained on a data mixture of the three human video datasets, as well as our in-house robot datasets, including Unitree G1, Fourier GR-1, AgiBot, and

SLIDE 5

Where it breaks

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 6

One-line takeaway

Trains multitask generalist robot policies through sequence modeling over heterogeneous human demonstrations.

Assembled from the paper's own PDF, parsed with its layout intact, 109,850 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Figures worth putting on a slide

  • Figure 1: **DreamDojo overview.** DreamDojo acquires comprehensive physical knowledge from large-scale human datasets by utilizing latent actions as unified labels. After post-trai
  • Figure 2: **Distribution analysis of DreamDojo-HV. (a)** Distribution of the scenarios and random examples from the most frequent categories. **(b)** [Left]: Distribution of subtas
  • Figure 3: **Latent action model.** [Left]: The information bottleneck design of our latent action model enforces action disentanglement, producing a continuous latent vector that r
  • Figure 4: **Benchmark visualization.** We rigorously construct six evaluation benchmarks that reflect the diverse scenarios and actions present in human datasets, while being out-o
  • Figure 5: **Downstream applications.** We show evidences that can be readily applied to benefit robot learning in policy evaluation without requiring real-world deployment, as well
  • Figure 6: **Live teleoperation.** We can teleoperate a virtual G1 robot using the PICO VR controller in real time.

Presented at

Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “Learning Generalist Robot Policies from Human Demonstrations”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Trains multitask generalist robot policies through sequence modeling over heterogeneous human demonstrations.

Read next

2024-10-15
Seonghyeon Ye, Joel Jang, Byeongguk Jeon +13 · 315 citations
2025-05-09
Qingwen Bu, Yanting Yang, Jisong Cai +5 · 400 citations
2025-06-11
Jiazhi Yang, Kashyap Chitta, Shenyuan Gao +7 · 47 citations

Something wrong on this page? Open a correction.