Robotics Papers

2026-06-29 · club pick

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

Yen-Jen Wang, Jiaman Li, Sirui Chen, Takara E. Truong, Pei Xu, Pieter Abbeel, Rocky Duan, Koushil Sreenath, Angjoo Kanazawa, Carmelo Sferrazza, Guanya Shi, Karen Liu

No peer-reviewed venue on record yet. 0 citations, as of the last refresh.

Abstract

Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale. We address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Our pipeline leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories using privileged scene information, and renders paired egocentric observations after the fact. We produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts short-horizon whole-body kinematic trajectories. A whole-body tracker converts these predictions into actions on the physical humanoid. We evaluate on the physical Unitree G1 performing navigation and single-object transport, demonstrating that synthesized interactions in reconstructed scenes provide effective supervision for sim-to-real perception-based humanoid loco-manipulation. Project Website: https://vision-language-kinematics.github.io/

arXiv comment: 19 pages, 7 figures, 4 tables

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Connecting egocentric observations to whole-body action is a fundamental challenge for humanoid robots operating in human-centered environments. A humanoid must perceive the world from its own viewpoint, identify task-relevant objects, navigate toward them, and physically interact with them through coordinated full-body motion. This perception-to-action loop, which humans execute effortlessly in everyday settings, requires a policy that maps high-dimensional visual input and language instructions to whole-body…

SLIDE 2

What came before

Vision-language and perception-based policies for humanoids. Vision-language-action (VLA) policies map visual observations and language instructions to robot actions and have become a promising paradigm for robot policy learning \[6, 7, [8\]](#page-8-0). Recent work extends this direction to humanoid robots, where policies must connect perception, language, and whole-body control \[9, 10, 11, 12, 13, 14, 15, 16,

SLIDE 3

The method

Our goal is to enable a Unitree G1 humanoid to perform perception-based loco-manipulation, including object-directed navigation and box interaction, from egocentric observations and task instructions. We formulate the problem as kinematic prediction followed by whole-body tracking: given the current egocentric view, instruction, and robot kinematic state, a high-level policy predicts a short-horizon G1 whole-body kinematic trajectory, and a whole-body tracker converts the predicted trajectory into executable robot…

SLIDE 4

What they measured

In this section, we introduce our hardware setup (Sec. 4.1) and then answer four research questions: (i) How efficiently can the proposed pipeline generate diverse paired VLK data? (Sec. 4.2) (ii) Can a VLK policy trained on synthesized data transfer to the physical humanoid? (Sec. 4.3\) (iii) How does the amount of synthesized training data affect performance across different tasks? (Sec. 4.4\) (iv) How does visual domain randomization help bridge the visual sim-to-real

SLIDE 5

Where it breaks

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 6

One-line takeaway

Video: https://youtu.be/ZB6k\iMJP7M

Assembled from the paper's own PDF, parsed with its layout intact, 68,994 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Figures worth putting on a slide

  • Figure 1: Synthetic interactions in reconstructed scenes enable real-world perception-based locomanipulation. Synthetic data generation produces paired egocentric observations, tas
  • Figure 2: Method overview. Top: we reconstruct 3D scenes, generate task waypoints, synthesize G1 motions, and render egocentric observations to produce paired vision-language-kinem
  • Figure 3: Examples of generated VLK supervision. Each sequence pairs an egocentric RGB observation and language instruction with the corresponding G1 whole-body kinematic trajector
  • Figure 3. For more examples, please see Appendix [A.6.](#page-14-0)
  • Figure 4: Ablation of training-data volume in the lab scene. We vary the number of synthesized trajectories per mode per layout and report success rates on the lab-scene validation
  • Figure 4. As shown in the figure, increasing the amount of synthesized training data consistently improves performance across the evaluated

Presented at

Saturday, August 1, 2026
Robotics & World Models Reading Club 21: Vision-Language-Kinematics Supervision for Perception-Based Humanoid Loco-Manipulation. SF 8/1
Listed on the event page as “Project: https://vision-language-kinematics.github.io”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Video: https://youtu.be/ZB6k\iMJP7M

Read next

Something wrong on this page? Open a correction.