Robotics Papers

2026-08-05 · IROS 2026

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

Jihoon Oh, Kento Kawaharazuka, Kei Okada

Published at IROS 2026. 0 citations, as of the last refresh.

Abstract

Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object-centric actionable affordances that are embodiment-agnostic. In this work, we propose a framework that leverages egocentric human videos with state-of-the-art 3D Structure-from-Motion and hand mesh reconstruction to extract actionable affordances such as visual, grasp, and trajectory affordances that explicitly encode where to interact, how to grasp, and how to move. We construct EgoAffordance, a large-scale dataset comprising 204K episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances. Building on this, we introduce VLAff, a large vision-language model-based unified foundation model that learns cross-modal correlations across all actionable affordances. Given a visual observation and instruction, VLAff generates visual affordance heatmaps, grasp poses, and trajectories, which are then converted into directly executable actions by utilizing 3D scene information. Through extensive experiments, we demonstrate that VLAff not only achieves state-of-the-art performance on visual affordance prediction, but can also be effectively applied to real robot applications such as zero-shot manipulation and affordance-guided robot learning.

arXiv comment: 8 pages, 5 figures. Accepted to IEEE/RSJ IROS 2026. Project page: https://ojh6404.github.io/vlaff/

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Foundation models trained on large-scale datasets have proven effective for many downstream tasks in natural language processing and computer vision, such as LLaMA and ChatGPT . In robotics, there have been attempts to build such foundation models. However, traditional paradigms like imitation learning face challenges: collecting robot action-state datasets is costly and timeconsuming, with task and environment domains often limited to laboratory

SLIDE 2

What came before

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 3

The method

Jihoon Oh<sup>1</sup> , Kento Kawaharazuka<sup>1</sup> , and Kei Okada<sup>1</sup> *Abstract*— Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn objectcentric actionable affordances that are embodiment-agnostic. In this work, we propose a framework that leverages egocentric human videos with state-of-the-art 3D Structure-from-Motion and hand mesh…

SLIDE 4

What they measured

We employ Qwen2.5-VL as our base vision-language model. For enhanced visual feature extraction, we integrate DINOv2 as our additional vision encoder. *b) Datasets:* Our final EgoAffordance dataset comprises 204,025 episodes, containing 5,782,431 visual heatmaps and 11,612,524 trajectory sequences. To improve fine-grained visual affordance prediction, we additionally incorporate data from HANDAL and SceneFun3D , which provide detailed object-level affordance annotations. *c) Visual Affordance:* We evaluate visual…

SLIDE 5

Where it breaks

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 6

One-line takeaway

This work proposes a framework that leverages egocentric human videos with state-of-the-art 3D Structure-from-Motion and hand mesh reconstruction to extract actionable affordances such as visual, grasp, and trajectory affordances that explicitly encode where to interact, how to grasp, and how to move.

Assembled from the paper's own PDF, parsed with its layout intact so tables and equations survive, 44,681 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.

Read next

2024-07-26
Mohan Kumar Srirama, Sudeep Dasari, Shikhar Bahl +1 · 50 citations
2025-09-24
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6 · 200 citations

Something wrong on this page? Open a correction.