Robotics Papers

2025-09-04 · RA-L · 35 citations · club pick

EMMA: Scaling Mobile Manipulation via Egocentric Human Data

Lawrence Y. Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, Danfei Xu

Published at RA-L (the arXiv record still lists it as a preprint). 35 citations, 2 of them influential, as of the last refresh.

Abstract

Scaling mobile manipulation imitation learning is bottlenecked by expensive mobile robot teleoperation. We present Egocentric Mobile MAnipulation (EMMA), an end-to-end framework training mobile manipulation policies from human mobile manipulation data with static robot data, sidestepping mobile teleoperation. To accomplish this, we co-train human full-body motion data with static robot data. In our experiments across three real-world tasks, EMMA demonstrates comparable performance to baselines trained on teleoperated mobile robot data (Mobile ALOHA), achieving higher or equivalent task performance in full task success. We find that EMMA is able to generalize to new spatial configurations and scenes, and we observe positive performance scaling as we increase the hours of human data, opening new avenues for scalable robotic learning in real-world environments. Details of this project can be found at https://ego-moma.github.io/.

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

MOBILE manipulation has emerged as one of the most challenging problems in robotics due to the dual demands of navigation and manipulation. While recent advances in robot policy learning have demonstrated impressive capabilities in static manipulation, extending these successes Manuscript received: August, 3, 2025; Revised November, 3, 2025; Accepted December, 12, 2025. The primary obstacle is data scarcity; current approaches tackling mobile manipulation rely on teleoperation frameworks akin to Mobile ALOHA…

SLIDE 2

What came before

Behavior Cloning (BC) has emerged as an effective approach for robot learning, where policies are trained with direct supervised learning from expert demonstrations. Recent advances have shown remarkable results [\[10\]](#page-7-9), [\[11\]](#page-7-10), [\[12\]](#page-7-11), [\[13\]](#page-7-12), [\[14\]](#page-7-13), [\[15\]](#page-7-14), including the promise of building general-purpose policies by learning from large-scale datasets [\[11\]](#page-7-10), [\[15\]](#page-7-14), [\[16\]](#page-7-15). In…

SLIDE 3

The method

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 4

What they measured

H1: EMMA can achieve performance comparable to systems trained on teleoperated mobile manipulation data. H2: Key design decisions of EMMA improve downstream task performance and robustness. H3: Given an initial amount of static robot manipulation data, it is more valuable to collect additional human mobile manipulation over mobile robot teleoperation data. We evaluate these hypotheses through four long horizon mobile manipulation tasks (Fig. 4\). *Table Service.* Two tables are set 2m

SLIDE 5

Where it breaks

Despite the strengths of EMMA, our approach inherits several inherent limitations from learning mobile manipulation primarily through human demonstrations. First, the framework assumes that the visual and spatial distributions encountered during robot deployment lie within, or close to, those seen in human demonstrations. This assumption may break down when the kinematic or viewpoint differences between human and robot embodiments become

SLIDE 6

One-line takeaway

Bridges humans and robots through shared latent action representations using contrastive alignment and cycle-consistency objectives.

Assembled from the paper's own PDF, parsed with its layout intact, 46,737 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Presented at

Saturday, May 16, 2026
Robotics & World Models Reading Club 08: Embodied Human Data as the “Internet of Motion and Behavior” — San Francisco 0516
Listed on the event page as “Cross-Embodiment Policy Learning via Representation Alignment”. The arXiv title above is the record.
San Francisco, CA
Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
San Francisco, CA
Why the club picked it. Bridges humans and robots through shared latent action representations using contrastive alignment and cycle-consistency objectives.

Read next

Something wrong on this page? Open a correction.