Robotics Papers

Trends

What the field is actually writing about

Every number here is a count arXiv itself returns, one query per topic per month, over 24 months of cs.RO. The query strings are printed next to the results so you can run them yourself and get the same answer.

Share of monthly cs.RO submissions

Share, not raw count. cs.RO itself is growing fast, so raw counts rise for everything. Share shows what is taking over.

0% 4% 9% 13% 17%Aug 24Dec 24Apr 25Aug 25Dec 25Apr 26Jul 26 Vision-language-actionWorld modelsTactile & force sensingFoundation modelsSimulationHumanoids & locomotion

Rising and fading

Change in share between the last six months and the twelve before them, in percentage points.

AreaPapers last monthShareChange24 months
Vision-language-action
Policies that consume pixels and language and emit robot actions directly, usually fine-tuned from a VLM backbone.
139 11.92% +7.1pp
World models
Learned simulators that predict future observations, so a policy can plan or train inside imagination instead of on hardware.
88 7.55% +2.85pp
Tactile & force sensing
Touch as a first-class modality: skins, current sensing, and contact-rich feedback loops.
78 6.69% +2.2pp
Foundation models & pretraining
Cross-embodiment pretraining, scaling laws, and the generalist-policy bet.
73 6.26% +1.9pp
Simulation, sim-to-real & benchmarks
Simulators, domain randomisation, and the evaluation harnesses that decide whether a policy really works.
70 6% +1.59pp
Humanoids & locomotion
Whole-body control, legged locomotion, and loco-manipulation on human-form robots.
69 5.92% +1.31pp
Data collection & teleoperation
The unglamorous bottleneck: rigs, interfaces, and pipelines that produce robot demonstrations.
79 6.78% +0.91pp
Egocentric & human data
Turning first-person human video and wearable capture into robot training signal, bypassing teleoperation.
53 4.55% +0.85pp
Dexterous manipulation
Multi-fingered hands, in-hand reorientation, and contact-rich control of everyday objects.
26 2.23% +0.75pp
Video & generative modeling
The upstream stack robotics borrows: video generation, diffusion and autoregressive models of pixels, and the representations they learn.
85 7.29% +0.73pp
Safety & evaluation
Does it fail safely, and how would we know? Verification, red-teaming, and honest metrics.
51 4.37% +0.63pp
Hardware & morphology co-design
Designing the body and the learning algorithm together instead of in sequence.
42 3.6% +0.48pp
3D & spatial reasoning
Reconstruction, scene representation, and the geometry a policy needs to act in a room it has never seen.
38 3.26% -0.16pp
Imitation & diffusion policies
Behaviour cloning at scale, and the generative-model action heads that made it work.
73 6.26% -0.32pp
Human-robot interaction
Robots among people: social robots, collaboration, assistance, and the interfaces between human intent and robot action.
53 4.55% -0.4pp
Navigation & mobility
Getting a body from A to B: mobile manipulation, exploration, and outdoor autonomy.
226 19.38% -0.83pp
Reinforcement learning & control
Online and offline RL, model-predictive control, and the safety filters wrapped around them.
179 15.35% -1.44pp

Organisations in the frame

Two signals, deliberately kept apart. In text counts papers whose title, abstract or arXiv comment names the organisation, which includes papers about someone else's robot. In author blocks counts the 31 papers we hold full text for and reads the affiliations directly. Neither is an affiliation database, and the second is a small sample.

OrganisationIn textIn author blocksTotal
Unitree 633 66
NVIDIA 487 55
Google DeepMind 186 24
Physical Intelligence 152 17
Stanford 49 13
Georgia Tech 17 8
Meta FAIR 62 8
UC Berkeley 17 8
AgiBot 6· 6
MIT 31 4
CMU 13 4
ByteDance 21 3
Toyota Research Institute 2· 2
ETH Zurich ·2 2
Boston Dynamics 11 2
Columbia ·1 1
UCSD ·1 1
Shanghai AI Lab 1· 1
Amazon ·1 1
Applied Intuition 1· 1
Princeton 1· 1

Vocabulary of the moment

The most common two-word and three-word technical phrases across 2,209 indexed abstracts. This is a description of the corpus on this site, not a measurement of the field. Use the share chart above for that.

vision-language-action vla 140dexterous manipulation 94action generation 61gaussian splatting 57autonomous driving 55vla policies 48motion planning 46human demonstrations 46sim-to-real transfer 44percentage points 43contact-rich manipulation 41visual observations 40manipulation policies 37control barrier 36predictive control 36page https 35real-world deployment 34action chunks 33vision-language-action vla policies 33tactile sensing 32project website 32human-robot interaction 31physical interaction 31latent space 30embodied agents 29challenging due 29action space 29diffusion policy 29website https 29dexterous hand 28human videos 28splatting dgs 26gaussian splatting dgs 26action prediction 26action expert 25unified framework 25neural network 25publicly available 25long-horizon manipulation 24often rely 24embodied intelligence 24dynamic environments 24real-world manipulation 24obstacle avoidance 23

How to check this

topic series

One arXiv API query per (topic, month) reading opensearch:totalResults for cat:cs.RO restricted to that month. Share is topic count / all cs.RO submissions that month. Queries are printed next to each topic so anyone can re-run them.

keywords

Bigrams and trigrams from titles and abstracts of the papers held in this corpus. Lift compares the term rate in the newest year against all earlier years in the same corpus. Small corpus, so treat as a pointer, not a measurement.

orgs

Two independent signals kept separate: (1) organisation names appearing in paper text or arXiv comments, (2) organisation names appearing in the author block of the papers we hold full text for. Neither is a complete affiliation database.

Generated 2026-08-28. Download the raw numbers.