Robotics Papers

2025-12-31 · 30 citations · club pick

Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow

Karthik Dharmarajan, Wenlong Huang, Jiajun Wu, Li Fei-Fei, Ruohan Zhang

No peer-reviewed venue on record yet. 30 citations, 6 of them influential, as of the last refresh.

Abstract

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the low-level actions demanded by robotic systems. We observe that given an initial image and task instruction, these models excel at synthesizing sensible object motions. Thus, we introduce Dream2Flow, a framework that bridges video generation and robotic control through 3D object flow as an intermediate representation. Our method reconstructs 3D object motions from generated videos and formulates manipulation as object trajectory tracking. By separating the state changes from the actuators that realize those changes, Dream2Flow overcomes the embodiment gap and enables zero-shot guidance from pre-trained video models to manipulate objects of diverse categories-including rigid, articulated, deformable, and granular. Through trajectory optimization or reinforcement learning, Dream2Flow converts reconstructed 3D object flow into executable low-level commands without task-specific demonstrations. Simulation and real-world experiments highlight 3D object flow as a general and scalable interface for adapting video generation models to open-world robotic manipulation. Videos and visualizations are available at https://dream2flow.github.io/.

arXiv comment: Project website: https://dream2flow.github.io/

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Robotic manipulation in the open world could greatly benefit from visual world models that predict how an environment would evolve given an agent's interactions. Recent advances in generative video modeling have produced systems capable of zero-shot synthesizing minute-long, highfidelity clips of physical interactions in pixel space, conditioned on an unseen initial image and an open-ended task instruction . Such video models implicitly capture intuitive physics and rich priors of object properties and…

SLIDE 2

What came before

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 3

The method

Recent work increasingly integrates video models across robotic tasks in various ways . They can serve as auxiliary training objectives [59–64], as reward models [65– 67], as policies , or as a simulator for the environments . Notably, predictive modeling in robotics can leverage video frame prediction as a form of world model. By simulating future visual observations, these models can enable visual planning and manipulation by anticipating how the environment will

SLIDE 4

What they measured

We seek to answer the following research questions through our experiments: Q1: What are properties of 3D object flow when used as an interface to bridge videos and robot control? Q2: How does Dream2Flow perform compared to alternative interfaces? Q3: How effective is 3D object flow as a reward for learning sensimotor policies? Q4: How does the choice of video model affect Dream2Flow in simulation and in real-world

SLIDE 5

Where it breaks

First, it relies on a rigid-grasp assumption for real world manipulation, limiting the types of tasks that can be performed. While this work shows that a particle dynamics model can be used for other types of tasks such as non-prehensile pushing, training and scaling a particle dynamics model for the real world is non-trivial and can be considered for future work. Another limitation is that the total processing time to get 3D object flow depending on the video generation model is between 3 and 11 minutes, which…

SLIDE 6

One-line takeaway

This work introduces Dream2Flow, a framework that bridges video generation and robotic control through 3D object flow as an intermediate representation and enables zero-shot guidance from pre-trained video models to manipulate objects of diverse categories-including rigid, articulated, deformable, and granular.

Assembled from the paper's own PDF, parsed with its layout intact, 70,511 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.

Presented at

Saturday, August 29, 2026
Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap+Video-Tactile-Action Model. SF 8/29
San Francisco, CA

Read next

2024-03-05
Patrick Esser, Sumith Kulal, Andreas Blattmann +14 · 4,594 citations
2025-01-23
Jie Liu, Gongye Liu, Jiajun Liang +14 · 248 citations
2026-07-21
Hadi Alzayer, Wenlong Huang, Haonan Chen +8 · 1 citation

Something wrong on this page? Open a correction.