Robotics Papers

2025-05-24 · 21 citations · club pick

ManiFeel: Benchmarking and Understanding Visuotactile Manipulation Policy Learning

Quan Khanh Luu, Pokuang Zhou, Zhengtong Xu, Zhiyuan Zhang, Qiang Qiu, Yu She

No peer-reviewed venue on record yet. 21 citations, as of the last refresh.

Abstract

Supervised visuomotor policies have shown strong performance in robotic manipulation but often struggle in tasks with limited visual inputs, such as operations in confined spaces and dimly lit environments, or tasks requiring precise perception of object properties and environmental interactions. In such cases, tactile feedback becomes essential for manipulation. While the rapid progress of supervised visuomotor policies has benefited greatly from high-quality, reproducible simulation benchmarks in visual imitation, the visuotactile domain still lacks a similarly comprehensive and reliable benchmark for large-scale and rigorous evaluation. To address this, we introduce ManiFeel, a reproducible and scalable simulation benchmark designed to systematically study supervised visuotactile policy learning. ManiFeel offers a diverse suite of contact-rich and visually challenging manipulation tasks, a modular evaluation pipeline spanning sensing modalities, tactile representations, and policy architectures, as well as real-world validation. Through extensive experiments, ManiFeel demonstrates how tactile sensing enhances policy performance across diverse manipulation scenarios, ranging from precise contact-driven operations to visually constrained settings. In addition, the results reveal task-dependent strengths of different tactile modalities and identify key design principles and open challenges for robust visuotactile policy learning. Real-world evaluations further confirm that ManiFeel provides a reliable and meaningful foundation for benchmarking and future visuotactile policy development. To foster reproducibility and future research, we will release our codebase, datasets, training logs, and pretrained checkpoints, aiming to accelerate progress toward generalizable visuotactile policy learning and manipulation.

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

In recent years, supervised policy learning has made significant strides in robot manipulation, where visual input is used to generate actions for long-horizon and dexterous manipulation tasks [\[1\]](#page-16-0)–[\[4\]](#page-16-1). However, vision-based policies face notable limitations not only in environments where visual cues are absent or severely degraded, but also in manipulation scenarios that inherently demand precise contact interactions. In cluttered or confined spaces, under low-light conditions, or…

SLIDE 2

What came before

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 3

The method

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 4

What they measured

This section introduces the task suite and modular pipeline design of the ManiFeel benchmark. ManiFeel features a diverse set of manipulation tasks spanning clear-vision conditions to severely degraded or fully occluded visual settings (Section III-A\). The benchmark includes three main task categories: *Insertion*, *Screwing*, and *Exploration*. This results in a total of 13 different task setups, including 9 simulation setups and 4 real-world setups, as shown in Fig.

SLIDE 5

Where it breaks

Our results yield several key insights that directly address the research questions posed in this study. First, across both simulation and real-world experiments, tactile sensing consistently improves policy performance in contact-rich manipulation and in scenarios where visual input is limited or ambiguous. These results confirm that tactile sensing provides complementary feedback that is essential for robust control in contact-dominated or visually uncertain

SLIDE 6

One-line takeaway

Website: https://zhengtongxu.github.io/manifeel-website/

Assembled from the paper's own PDF, parsed with its layout intact, 110,550 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Presented at

Saturday, July 11, 2026
Robotics & World Models Reading Club 17: Soft Tactile-Centric Multimodal Intelligence Toward Safe and Dexterous Manipulation. SF 07/11
Listed on the event page as “ManiFeel preprint and project website:”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Website: https://zhengtongxu.github.io/manifeel-website/

Read next

Something wrong on this page? Open a correction.