Robotics Papers

2026-08-18 · ECCV 2026

Reproducible Multimodal Affordance Prediction

Tommaso Apicella, Alessio Xompero, Andrea Cavallaro

Published at ECCV 2026. 0 citations, as of the last refresh.

Abstract

Affordance prediction is the identification of potential actions an agent can perform on a target object from multimodal inputs. Affordance prediction methods are difficult to evaluate and compare due to heterogeneous problem formulations, inconsistent dataset annotations, incomplete reporting of experimental protocols, and limited information about deployment conditions. These limitations challenge fair benchmarking and performance comparison. To promote transparency, we propose the Affordance Sheet, a documentation detailing task formulation with its input modalities, model architectures and training information, datasets, and experimental protocols. Affordance Sheets enable reproducible benchmarking and reliable evaluation of affordance models for real-world scenarios, including generalisation to novel conditions and human safety.

arXiv comment: Paper accepted to Workshop on Human-Centered Multimodal Intelligence in the Wild (HCMIW) in European Conference on Computer Vision (ECCV) 2026; 18 pages, 3 figures, 7 tables. Project webpage at https://apicis.github.io/aff-sheet

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Affordances describe the potential actions that an agent can perform on objects in the environment [\[20\]](#page-15-0), inferred from multimodal sensory observations such as visual, depth, acoustic, and tactile data [\[28\]](#page-15-1). Understanding affordances enables the agent to accomplish a task, selecting which objects in the environment to interact with, what actions to perform, and how to execute them. This reasoning within and across modalities, beyond simply perceiving scene and objects, supports…

SLIDE 2

What came before

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 3

The method

The lack of publicly available implementation of methods (RC2) \[22,56–[58\]](#page-17-7), the lack of publicly available trained models (RC3) \[22, 42, 56[–58\]](#page-17-7), and the lack of details of experimental setups (RC4) \[14,22,42,56[–58\]](#page-17-7) can challenge researchers in reproducing previous works for comparative evaluations. The release of the model trained weights and the implementation of the method and inference pipeline is a crucial aspect for reproducibility, especially for deep-learning…

SLIDE 4

What they measured

Table 1 reports the main characteristics of the available datasets and benchmarks in affordance prediction. None of the datasets is collected for benchmarking methods under different in-the-wild conditions (RC1), such as illumination, clutter, or hand-occlusions. Each affordance dataset is collected for a specific formulation, and not re-used across different affordance tasks, limiting fair comparison and preventing comprehensive benchmarking. One of the main obstacles to obtain large-scale data collections to…

SLIDE 5

Where it breaks

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 6

One-line takeaway

The Affordance Sheet is proposed, a documentation detailing task formulation with its input modalities, model architectures and training information, datasets, and experimental protocols, which enable reproducible benchmarking and reliable evaluation of affordance models for real-world scenarios, including generalisation to novel conditions and human safety.

Assembled from the paper's own PDF, parsed with its layout intact so tables and equations survive, 85,562 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.

Read next

Something wrong on this page? Open a correction.