2026-06-29 · club pick
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
Yen-Jen Wang, Jiaman Li, Sirui Chen, Takara E. Truong, Pei Xu, Pieter Abbeel, Rocky Duan, Koushil Sreenath, Angjoo Kanazawa, Carmelo Sferrazza, Guanya Shi, Karen Liu
No peer-reviewed venue on record yet. 0 citations, as of the last refresh.
Abstract
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale. We address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Our pipeline leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories using privileged scene information, and renders paired egocentric observations after the fact. We produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts short-horizon whole-body kinematic trajectories. A whole-body tracker converts these predictions into actions on the physical humanoid. We evaluate on the physical Unitree G1 performing navigation and single-object transport, demonstrating that synthesized interactions in reconstructed scenes provide effective supervision for sim-to-real perception-based humanoid loco-manipulation. Project Website: https://vision-language-kinematics.github.io/
arXiv comment: 19 pages, 7 figures, 4 tables
Ten-minute slide kit
Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.
Assembled from the paper's own PDF, parsed with its layout intact, 68,994 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.
Figures worth putting on a slide
- Figure 1: Synthetic interactions in reconstructed scenes enable real-world perception-based locomanipulation. Synthetic data generation produces paired egocentric observations, tas
- Figure 2: Method overview. Top: we reconstruct 3D scenes, generate task waypoints, synthesize G1 motions, and render egocentric observations to produce paired vision-language-kinem
- Figure 3: Examples of generated VLK supervision. Each sequence pairs an egocentric RGB observation and language instruction with the corresponding G1 whole-body kinematic trajector
- Figure 3. For more examples, please see Appendix [A.6.](#page-14-0)
- Figure 4: Ablation of training-data volume in the lab scene. We vary the number of synthesized trajectories per mode per layout and report success rates on the lab-scene validation
- Figure 4. As shown in the figure, increasing the amount of synthesized training data consistently improves performance across the evaluated
Presented at
Read next
Something wrong on this page? Open a correction.