2025-11-20 · 16 citations · club pick
Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
Irmak Guzey, Haozhi Qi, Julen Urain, Changhao Wang, Jessica Yin, Krishna Bodduluri, Mike Lambeta, Lerrel Pinto, Akshara Rai, Jitendra Malik, Tingfan Wu, Akash Sharma, Homanga Bharadhwaj
No peer-reviewed venue on record yet. 16 citations, as of the last refresh.
Abstract
Learning multi-fingered robot policies from humans performing daily tasks in natural environments has long been a grand goal in the robotics community. Achieving this would mark significant progress toward generalizable robot manipulation in human environments, as it would reduce the reliance on labor-intensive robot data collection. Despite substantial efforts, progress toward this goal has been bottle-necked by the embodiment gap between humans and robots, as well as by difficulties in extracting relevant contextual and motion cues that enable learning of autonomous policies from in-the-wild human videos. We claim that with simple yet sufficiently powerful hardware for obtaining human data and our proposed framework AINA, we are now one significant step closer to achieving this dream. AINA enables learning multi-fingered policies from data collected by anyone, anywhere, and in any environment using Aria Gen 2 glasses. These glasses are lightweight and portable, feature a high-resolution RGB camera, provide accurate on-board 3D head and hand poses, and offer a wide stereo view that can be leveraged for depth estimation of the scene. This setup enables the learning of 3D point-based policies for multi-fingered hands that are robust to background changes and can be deployed directly without requiring any robot data (including online corrections, reinforcement learning, or simulation). We compare our framework against prior human-to-robot policy learning approaches, ablate our design choices, and demonstrate results across nine everyday manipulation tasks. Robot rollouts are best viewed on our website: https://aina-robot.github.io.
Ten-minute slide kit
Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.
SLIDE 1The problem
*"The most profound technologies are those that disappear. They weave themselves into the fabric of everyday life until they are indistinguishable from it."* *— Mark Weiser, 1991* Robots autonomously performing diverse manipulation tasks by watching humans go about their daily lives has been a dream in Artificial Intelligence (AI) for decades. However, this remains challenging due to the embodiment gap between humans and robots, as well as the disparity between human video views and the sensor perspectives of a…
SLIDE 2What came before
Our aim is to develop a *simple* framework for closedloop policy learning capable of performing diverse everyday manipulation tasks with dexterous multi-fingered hands. These <span id="page-2-1"></span> Fig. 4: Illustration of our overall AINA framework. On the left, we show how the data is processed: the human hand pose is extracted directly by the Aria Gen 2 glasses, and stereo depth is estimated from the surrounding SLAM camera
SLIDE 3The method
Not recoverable from the parsed text. Read this section in the paper yourself.
SLIDE 4What they measured
These are collected with natural human motions and with the right hand performing the respective tasks (no additional sensors on the humans or the environments, except Aria glasses). - operation space changes? - 4) How well does AINA generalize spatially and across different
SLIDE 5Where it breaks
In this work, we presented AINA, a new framework that leverages capabilities of Aria Gen 2 glasses to learn pointbased multi-fingered policies from explicitly in-the-wild human demonstrations. While promising, we observe three limitations. First, our framework cannot easily integrate force feedback, since hand pose estimation alone cannot capture this information, which is often crucial for accurate dexterous manipulation \[57, 58,
SLIDE 6One-line takeaway
The proposed framework AINA enables the learning of 3D point-based policies for multi-fingered hands that are robust to background changes and can be deployed directly without requiring any robot data (including online corrections, reinforcement learning, or simulation).
Assembled from the paper's own PDF, parsed with its layout intact, 55,188 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.
Presented at
Saturday, June 20, 2026
Robotics & World Models Reading Club 13: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
San Francisco, CA
Read next
2025-05-26
Vincent Liu, Ademi Adeniji, Haotian Zhan +4 · 41 citations
2023-09-21
Irmak Guzey, Yinlong Dai, Ben Evans +2 · 60 citations
2024-04-25
Toru Lin, Yu Zhang, Qiyang Li +4 · 148 citations
2026-06-17
Bhawna Paliwal, Haritheja Etukuru, William Liang +3 · 5 citations
2022-10-10
Haozhi Qi, Ashish Kumar, Roberto Calandra +2 · 203 citations