Robotics Papers

2026-03-24 · 12 citations · club pick

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Haoran Yuan, Weigang Yi, Zhenyu Zhang, Wendi Chen, Yuchen Mo, Jiashi Yin, Xinzhuo Li, Xiangyu Zeng, Chuan Wen, Cewu Lu, Katherine Driggs-Campbell, Ismini Lourentzou

No peer-reviewed venue on record yet. 12 citations, as of the last refresh.

Abstract

Video-Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent action predictions. Although such models demonstrate strong performance on long-horizon tasks through visual reasoning, they remain limited in contact-rich scenarios where critical interaction states are only partially observable from vision alone. In particular, fine-grained force modulation and contact transitions are not reliably encoded in visual tokens, leading to unstable or imprecise behaviors. To bridge this gap, we introduce the Video-Tactile Action Model (VTAM), a multimodal world modeling framework that incorporates tactile perception as a complementary grounding signal. VTAM augments a pretrained video transformer with tactile streams via a lightweight modality transfer finetuning, enabling efficient cross-modal representation learning without tactile-language paired data or independent tactile pretraining. To stabilize multimodal fusion, we introduce a tactile regularization loss that enforces balanced cross-modal attention, preventing visual latent dominance in the action model. VTAM demonstrates superior performance in contact-rich manipulation, maintaining a robust success rate of 90 percent on average. In challenging scenarios such as potato chip pick-and-place requiring high-fidelity force awareness, VTAM outperforms the pi 0.5 baseline by 80 percent. Our findings demonstrate that integrating tactile feedback is essential for correcting visual estimation errors in world action models, providing a scalable approach to physically grounded embodied foundation models.

arXiv comment: https://plan-lab.github.io/projects/vtam/

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Recent advances in Vision–Language–Action (VLA) models have enabled generalist robot control through large-scale multimodal alignment \[3, 17, [49\]](#page-15-0). By embedding visual observations and language instructions into a shared semantic latent space, these models can generalize across diverse manipulation tasks and environments \[9, [30\]](#page-14-0). However, while vision supports high-level semantic understanding and language specifies task intent, *physical interaction is fundamentally governed by…

SLIDE 2

What came before

**Vision-Language-Action Models.** VLA models have emerged as the dominant paradigm for generalist robot control, leveraging internet-scale vision–language pretraining to ground natural-language instructions in visual observations and decode motor commands through a unified architecture \[3, 4, 17, 30, [49\]](#page-15-0). Subsequent efforts have expanded the paradigm along several axes, incorporating 3D geometric priors \[28, [46\]](#page-15-1), hierarchical task planning \[1, [23\]](#page-13-4), and predictive…

SLIDE 3

The method

We present the Video-Tactile Action Model (VTAM), a unified visuo-tactile world action model designed for contact-rich manipulation. As illustrated in Figure 2, VTAM operates by projecting both multi-view visual observations and high-resolution tactile streams (e.g., GelSight [\[41\]](#page-14-11)) into a shared continuous latent space via a pre-trained Variational Autoencoder (VAE). Within this space, a multi-view diffusion process employing alternating intra-view and cross-view attention jointly models the…

SLIDE 4

What they measured

We evaluate VTAM on real-world contact-rich manipulation tasks to study the effectiveness of visuo–tactile world action modeling. Our experiments aim to answer the following key questions: - **Q1: Effectiveness of Visuo-Tactile World Action Modeling.** Does VTAM outperform vision-only and multimodal baselines in scenarios requiring fine-grained force modulation? - **Q2: Latent Video Fusion vs. Late-stage Injection.** Does modeling visuo-tactile dynamics within a shared video latent space offer performance…

SLIDE 5

Where it breaks

Figure 4 shows qualitative comparisons across the three tasks. We analyze the behaviors of different methods to understand how VTAM addresses the challenges in contact-rich tasks. Additional examples can be found in the Appendix. **Chip Pick-and-Place.** For the GE vision-only baseline, the main failure arises from the inability to verify successful

SLIDE 6

One-line takeaway

This work introduces the Video-Tactile Action Model (VTAM), a multimodal world modeling framework that incorporates tactile perception as a complementary grounding signal, and demonstrates superior performance in contact-rich manipulation.

Assembled from the paper's own PDF, parsed with its layout intact, 61,478 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.

Figures worth putting on a slide

  • Figure 1:** We introduce **VTAM**, a generalist **V**ideo–**T**actile **A**ction **M**odel that integrates tactile sensing into a predictive video world model. By grounding control
  • Figure 2: VTAM Overview.** A pretrained video backbone jointly models multi-view visual and tactile latents via alternating intra-view and cross-view attention. The resulting multi
  • Figure 3: Experiment setup and data acquisition.** We collect demonstrations through manual teleoperation using a visuo–tactile sensing setup for contact-rich manipulation tasks su
  • Figure 4: Qualitative comparison between VTAM and baseline methods on real-world manipulation tasks.** Top: *Chip pick-and-place*. Vision-only baselines fail to determine whether t
  • Figure 5: Prediction visualization of the backbone video model.** From top to bottom: Camera-1 view, Camera-2 view, Tactile stream prediction. Ground-truth (top rows) and VTAM pred
  • Figure 6: Qualitative peeling results.** VTAM achieves an 85% success rate (17/20 trials), producing peel strips longer than 10 cm in successful runs.

Presented at

Saturday, August 29, 2026
Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap+Video-Tactile-Action Model. SF 8/29
San Francisco, CA

Read next

Something wrong on this page? Open a correction.