Robotics Papers

2025-05-28 · 111 citations · club pick

DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation

Mengda Xu, Han Zhang, Yifan Hou, Zhenjia Xu, Linxi Fan, Manuela Veloso, Shuran Song

No peer-reviewed venue on record yet. 111 citations, 5 of them influential, as of the last refresh.

Abstract

We present DexUMI - a data collection and policy learning framework that uses the human hand as the natural interface to transfer dexterous manipulation skills to various robot hands. DexUMI includes hardware and software adaptations to minimize the embodiment gap between the human hand and various robot hands. The hardware adaptation bridges the kinematics gap using a wearable hand exoskeleton. It allows direct haptic feedback in manipulation data collection and adapts human motion to feasible robot hand motion. The software adaptation bridges the visual gap by replacing the human hand in video data with high-fidelity robot hand inpainting. We demonstrate DexUMI's capabilities through comprehensive real-world experiments on two different dexterous robot hand hardware platforms, achieving an average task success rate of 86%.

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Human hands are incredibly dexterous in a wide range of tasks. Dexterous robot hands are designed with the hope of replicating this capability. However, it remains a significant challenge to transfer skills from human hands to robotic counterparts due to their substantial *embodiment

SLIDE 2

What came before

Although extensive work has studied how to enable learning in simulated environments \[6–[20\]](#page-11-0), we focus on reviewing real world data collection methods. Teleoperation: Teleoperation is a popular interface for dexterous manipulation. Hand control is achieved with motion capture gloves [\[21–25\]](#page-11-0), virtual-reality devices \[26[–28\]](#page-12-0), or camera-based tracking

SLIDE 3

The method

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 4

What they measured

Target robot hands: We evaluate DexUMI across two different robot hands: - *Inspire Hand (IHand):* A twelve-DoF (six active DoFs) underactuated hand. The thumb contains three DoFs, the index finger has three DoFs, and each of the remaining fingers has two DoFs. Tasks: We evaluate DexUMI across four different real-world tasks: - Cube [IHand]: Pick up a 2.5cm wide cube from a table and place it into a cup. The main challenge is to stably operate the deformable tweezers with multifinger contacts. - Kitchen [XHand]:…

SLIDE 5

Where it breaks

We would like to discuss DexUMI's limitations from three different aspects: hardware adaptation, software adaptation, and existing robot hand

SLIDE 6

One-line takeaway

Learns view-invariant action representations by jointly leveraging egocentric and exocentric video data.

Assembled from the paper's own PDF, parsed with its layout intact, 75,104 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Figures worth putting on a slide

  • Figure 1: DexUMI transfer dexterous human manipulation skills to various robot hand by using wearable exoskeletons and a data processing framework. We demonstrate DexUMI's capabili
  • Figure 2: Exoskeleton Design. The optimized exoskeleton design shares the same joint-to-fingertip position mapping as the target robot hand while maintaining the wearability. The e
  • Figure 3: Mechanism Optimization. To avoid thumb collision between human hand and exoskeleton, the hardware optimization step allows us to move the exoskeleton thumb backward while
  • Figure 4: Bridging the Visual Gap. To convert the visual observation into policy training data, we first segment the exoskeleton using SAM2 (b) and inpaint the missing background (
  • Figure 5: Policy Rollout: We evaluate DexUMI's capabilities across challenging real-world tasks. The Cube task tests basic picking precision. The Egg Carton task evaluates multi-fi
  • Figure 6: Comparisons. a) The policy outputs relative hand actions yield more precise action and demonstrate better multi-finger coordination. Note, we draw a sketch for the knob c

Presented at

Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “Ego-Exo Transfer: Learning Action from First- and Third-Person Data”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Learns view-invariant action representations by jointly leveraging egocentric and exocentric video data.

Read next

Something wrong on this page? Open a correction.