Robotics Papers

2026-01-07 · 4 citations · club pick

Choreographing a World of Dynamic Objects

Yanzhe Lyu, Chen Geng, Karthik Dharmarajan, Yunzhi Zhang, Hadi Alzayer, Shangzhe Wu, Jiajun Wu

No peer-reviewed venue on record yet. 4 citations, 1 of them influential, as of the last refresh.

Abstract

Dynamic objects in our physical 4D (3D + time) world are constantly evolving, deforming, and interacting with other objects, leading to diverse 4D scene dynamics. In this paper, we present a universal generative pipeline, CHORD, for CHOReographing Dynamic objects and scenes and synthesizing this type of phenomena. Traditional rule-based graphics pipelines to create these dynamics are based on category-specific heuristics, yet are labor-intensive and not scalable. Recent learning-based methods typically demand large-scale datasets, which may not cover all object categories in interest. Our approach instead inherits the universality from the video generative models by proposing a distillation-based pipeline to extract the rich Lagrangian motion information hidden in the Eulerian representations of 2D videos. Our method is universal, versatile, and category-agnostic. We demonstrate its effectiveness by conducting experiments to generate a diverse range of multi-body 4D dynamics, show its advantage compared to existing methods, and demonstrate its applicability in generating robotics manipulation policies. Project page: https://yanzhelyu.github.io/chord

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Humans and other embodied agents live in a 4D (3D + time) world, a world composed of a diverse range of dynamic objects, *i.e*., objects that can evolve, deform, or interact with other objects. Creating 4D motions for both object deformations and interactions is crucial when building 3D world <sup>∗</sup>Equal contribution. †Work was done when Y. Lyu was a visiting student at Stanford

SLIDE 2

What came before

Generating 4D consistent object deformations has been a long-standing challenge in the community. Traditional approaches first determine category-specific kinematic models (*i.e*., rigging representations) \[6, 8, 22, 43, 45, 51, 71, [88\]](#page-11-2) and then generate motion based on them \[20, 28, 42, 47, 52, 58, 60, 63, [77\]](#page-11-3), which inherently limits these methods to constrained categories. Some methods \[73, 82, [86\]](#page-11-5) attempt to learn end-toend 4D generators from existing 4D object…

SLIDE 3

The method

Figure 2 shows an overview of our method. We iteratively optimize a 4D scene motion representation using guidance signals distilled from a video generative model. In the following section, we detail the three main components in this framework: a strategy for distillation from modern rectified flow-based video generative models (Sec. 3.2\), a robust and general 4D scene motion representation (Sec. 3.3\), and regularization terms to ensure stable optimization (Sec. 3.4\). The above-mentioned 4D SDS algorithm is…

SLIDE 4

What they measured

We evaluate our proposed method on a diverse dynamic scenes featuring multiple interacting

SLIDE 5

Where it breaks

Our failure cases mainly arise from two factors: (1) limitations of the underlying video generative model, and (2) the inability to handle objects that do not exist in the static snapshot but appear later in the motion sequence. Examples are shown in Our failure cases mainly arise from two factors: (1) limitations of the underlying video generative model, and (2) the inability to handle objects that do not exist in the static snapshot but appear later in the motion sequence. Examples are <span…

SLIDE 6

One-line takeaway

Keynote 2: VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Assembled from the paper's own PDF, parsed with its layout intact, 78,114 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Figures worth putting on a slide

  • Figure 1. 4D scene motion generated by our method. We present CHORD, a universal generative pipeline capable of animating scenes with multiple objects that interact with each other
  • Figure 2. Overview. For the input meshes of a given scene, we first convert them into 3D-GS representations to enable smooth gradient computation. The converted 3D-GS models are th
  • Figure 3. Illustration of the hierarchical control point representation. We represent the deformation using a spatial hierarchical structure. Coarse control points capture large-sc
  • Figure 4. Illustration of the Fenwick Tree representation. Each node stores the cumulative deformation over a temporal range, allowing nearby frames to share parameters and natural
  • Figure 5. Qualitative comparisons. We compare our approach with several mesh animation methods. Our method produces results that better align with the given prompts and exhibit mor
  • Figure 6. Real-world object animation results.

Presented at

Saturday, August 29, 2026
Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap+Video-Tactile-Action Model. SF 8/29
San Francisco, CA
Why the club picked it. Keynote 2: VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Read next

2019-06-10
RSS
2026-07-21
Hadi Alzayer, Wenlong Huang, Haonan Chen +8 · 1 citation

Something wrong on this page? Open a correction.