Robotics Papers

2026-08-13 · ECCV 2026

EgoPHI: Estimating Contact and Force from Egocentric Vision

Andela Ilic, Rachel Schuchert, Yijing Jiang, Christian Holz

Published at ECCV 2026. 0 citations, as of the last refresh.

Abstract

Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet reasoning about physically grounded interaction requires estimating the forces acting on hands and objects, beyond localizing contact. We present EgoPHI, the first method that jointly estimates dense contact maps and 3D force distributions on hand and object meshes from a single monocular RGB image and object geometry. To address the lack of scalable ground-truth force annotations, we introduce a physics-based simulation pipeline that augments existing hand-object datasets with dense per-vertex force supervision. EgoPHI then learns dense 3D contact and force on interacting hand and articulated object meshes, extending vision-based force estimation beyond image-space or planar settings. Our evaluation on in-distribution and out-of-distribution benchmarks shows that EgoPHI improves force estimation over existing approaches while generalizing to unseen datasets. To evaluate sim-to-real transfer, we constructed two physical objects that capture dense object contact and force magnitude and used them to record a dataset of interactions from eight participants across diverse touch and grasp types. Our results demonstrate that EgoPHI recovers meaningful 3D contact and force distributions in simulated, out-of-distribution, and real-world settings, advancing egocentric hand-object understanding from contact localization toward physically grounded interaction reasoning.

arXiv comment: Accepted by ECCV 2026

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

Understanding hand–object interaction is fundamental to modeling how humans engage with the physical world. Detecting contact and estimating the magnitude and spatial distribution of forces during interaction reveal intent [\[30\]](#page-16-0), stability, and control [\[20\]](#page-16-1) in object manipulation. For example, in robotics, recovering such information can enable robots to infer human manipulation strategies and learn stable grasping or object handling from

SLIDE 2

What came before

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 3

The method

Figure 2 presents an overview of the EgoPHI pipeline, which jointly estimates the object pose and per-vertex interaction forces and contacts. These force and <span id="page-4-0"></span> Fig. 2: EgoPHI's 3-stage pipeline: (1) visual & geometric feature extraction with crossmodal fusion, (2) object pose estimation, and (3) contact and force estimation. Our novel Graph-Based Interaction Blocks perform the core 3D reasoning by encoding intramesh structure and inter-mesh relationships between the hands and object.…

SLIDE 4

What they measured

Baselines We evaluate EgoPHI against existing methods, noting that prior work focuses exclusively on the hand; to our knowledge, no existing model estimates contact or force on the manipulated object. While EgoPressure [\[47\]](#page-17-4) addresses 3D hand pressure, its code is not publicly available. Therefore, we establish two baseline strategies. 3D Mesh-Based Comparison: We utilize the state-of-the-art contact estimation model HACO [\[21\]](#page-16-3). For contact, we use the publicly available version…

SLIDE 5

Where it breaks

Our evaluation shows that egocentric vision, when combined with known object geometry, can effectively infer dense 3D force fields. EgoPHI generalizes from simulated supervision to unseen interactions and, to some extent, real-world objects, showing that visual cues support physically meaningful force estimation. First, EgoPHI processes frames independently and discards temporal

SLIDE 6

One-line takeaway

This work presents EgoPHI, the first method that jointly estimates dense contact maps and 3D force distributions on hand and object meshes from a single monocular RGB image and object geometry, and demonstrates that EgoPHI improves force estimation over existing approaches while generalizing to unseen datasets.

Assembled from the paper's own PDF, parsed with its layout intact so tables and equations survive, 54,530 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. Check it before you present it.

Read next

Something wrong on this page? Open a correction.