2026-03-24 · 12 citations · club pick
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
Haoran Yuan, Weigang Yi, Zhenyu Zhang, Wendi Chen, Yuchen Mo, Jiashi Yin, Xinzhuo Li, Xiangyu Zeng, Chuan Wen, Cewu Lu, Katherine Driggs-Campbell, Ismini Lourentzou
No peer-reviewed venue on record yet. 12 citations, as of the last refresh.
Abstract
Video-Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent action predictions. Although such models demonstrate strong performance on long-horizon tasks through visual reasoning, they remain limited in contact-rich scenarios where critical interaction states are only partially observable from vision alone. In particular, fine-grained force modulation and contact transitions are not reliably encoded in visual tokens, leading to unstable or imprecise behaviors. To bridge this gap, we introduce the Video-Tactile Action Model (VTAM), a multimodal world modeling framework that incorporates tactile perception as a complementary grounding signal. VTAM augments a pretrained video transformer with tactile streams via a lightweight modality transfer finetuning, enabling efficient cross-modal representation learning without tactile-language paired data or independent tactile pretraining. To stabilize multimodal fusion, we introduce a tactile regularization loss that enforces balanced cross-modal attention, preventing visual latent dominance in the action model. VTAM demonstrates superior performance in contact-rich manipulation, maintaining a robust success rate of 90 percent on average. In challenging scenarios such as potato chip pick-and-place requiring high-fidelity force awareness, VTAM outperforms the pi 0.5 baseline by 80 percent. Our findings demonstrate that integrating tactile feedback is essential for correcting visual estimation errors in world action models, providing a scalable approach to physically grounded embodied foundation models.
arXiv comment: https://plan-lab.github.io/projects/vtam/
Ten-minute slide kit
Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.
Assembled from the paper's own PDF, parsed with its layout intact, 61,478 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.
Figures worth putting on a slide
- Figure 1:** We introduce **VTAM**, a generalist **V**ideo–**T**actile **A**ction **M**odel that integrates tactile sensing into a predictive video world model. By grounding control
- Figure 2: VTAM Overview.** A pretrained video backbone jointly models multi-view visual and tactile latents via alternating intra-view and cross-view attention. The resulting multi
- Figure 3: Experiment setup and data acquisition.** We collect demonstrations through manual teleoperation using a visuo–tactile sensing setup for contact-rich manipulation tasks su
- Figure 4: Qualitative comparison between VTAM and baseline methods on real-world manipulation tasks.** Top: *Chip pick-and-place*. Vision-only baselines fail to determine whether t
- Figure 5: Prediction visualization of the backbone video model.** From top to bottom: Camera-1 view, Camera-2 view, Tactile stream prediction. Ground-truth (top rows) and VTAM pred
- Figure 6: Qualitative peeling results.** VTAM achieves an 85% success rate (17/20 trials), producing peel strips longer than 10 cm in successful runs.
Presented at
Read next
Something wrong on this page? Open a correction.