2025-12-31 · 30 citations · club pick
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
Karthik Dharmarajan, Wenlong Huang, Jiajun Wu, Li Fei-Fei, Ruohan Zhang
No peer-reviewed venue on record yet. 30 citations, 6 of them influential, as of the last refresh.
Abstract
Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the low-level actions demanded by robotic systems. We observe that given an initial image and task instruction, these models excel at synthesizing sensible object motions. Thus, we introduce Dream2Flow, a framework that bridges video generation and robotic control through 3D object flow as an intermediate representation. Our method reconstructs 3D object motions from generated videos and formulates manipulation as object trajectory tracking. By separating the state changes from the actuators that realize those changes, Dream2Flow overcomes the embodiment gap and enables zero-shot guidance from pre-trained video models to manipulate objects of diverse categories-including rigid, articulated, deformable, and granular. Through trajectory optimization or reinforcement learning, Dream2Flow converts reconstructed 3D object flow into executable low-level commands without task-specific demonstrations. Simulation and real-world experiments highlight 3D object flow as a general and scalable interface for adapting video generation models to open-world robotic manipulation. Videos and visualizations are available at https://dream2flow.github.io/.
arXiv comment: Project website: https://dream2flow.github.io/
Ten-minute slide kit
Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.
SLIDE 1The problem
Robotic manipulation in the open world could greatly benefit from visual world models that predict how an environment would evolve given an agent's interactions. Recent advances in generative video modeling have produced systems capable of zero-shot synthesizing minute-long, highfidelity clips of physical interactions in pixel space, conditioned on an unseen initial image and an open-ended task instruction . Such video models implicitly capture intuitive physics and rich priors of object properties and…
SLIDE 2What came before
Not recoverable from the parsed text. Read this section in the paper yourself.
SLIDE 3The method
Recent work increasingly integrates video models across robotic tasks in various ways . They can serve as auxiliary training objectives [59–64], as reward models [65– 67], as policies , or as a simulator for the environments . Notably, predictive modeling in robotics can leverage video frame prediction as a form of world model. By simulating future visual observations, these models can enable visual planning and manipulation by anticipating how the environment will
SLIDE 4What they measured
We seek to answer the following research questions through our experiments: Q1: What are properties of 3D object flow when used as an interface to bridge videos and robot control? Q2: How does Dream2Flow perform compared to alternative interfaces? Q3: How effective is 3D object flow as a reward for learning sensimotor policies? Q4: How does the choice of video model affect Dream2Flow in simulation and in real-world
SLIDE 5Where it breaks
First, it relies on a rigid-grasp assumption for real world manipulation, limiting the types of tasks that can be performed. While this work shows that a particle dynamics model can be used for other types of tasks such as non-prehensile pushing, training and scaling a particle dynamics model for the real world is non-trivial and can be considered for future work. Another limitation is that the total processing time to get 3D object flow depending on the video generation model is between 3 and 11 minutes, which…
SLIDE 6One-line takeaway
This work introduces Dream2Flow, a framework that bridges video generation and robotic control through 3D object flow as an intermediate representation and enables zero-shot guidance from pre-trained video models to manipulate objects of diverse categories-including rigid, articulated, deformable, and granular.
Assembled from the paper's own PDF, parsed with its layout intact, 70,511 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.
Presented at
Saturday, August 29, 2026
Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap+Video-Tactile-Action Model. SF 8/29
San Francisco, CA
Read next
2024-03-05
Patrick Esser, Sumith Kulal, Andreas Blattmann +14 · 4,594 citations
2025-01-23
Jie Liu, Gongye Liu, Jiajun Liang +14 · 248 citations
2026-07-21
Hadi Alzayer, Wenlong Huang, Haonan Chen +8 · 1 citation
2025-10-09
Hongyu Li, Lingfeng Sun, Yafei Hu +4 · 49 citations
2023-12-21
Dan Kondratyuk, Lijun Yu, Xiuye Gu +28 · 544 citations