topic series
One arXiv API query per (topic, month) reading opensearch:totalResults for cat:cs.RO restricted to that month. Share is topic count / all cs.RO submissions that month. Queries are printed next to each topic so anyone can re-run them.
Trends
Every number here is a count arXiv itself returns, one query per topic per month, over 24 months of cs.RO. The query strings are printed next to the results so you can run them yourself and get the same answer.
Share, not raw count. cs.RO itself is growing fast, so raw counts rise for everything. Share shows what is taking over.
Change in share between the last six months and the twelve before them, in percentage points.
| Area | Papers last month | Share | Change | 24 months |
|---|---|---|---|---|
| Vision-language-action Policies that consume pixels and language and emit robot actions directly, usually fine-tuned from a VLM backbone. |
139 | 11.92% | +7.1pp | |
| World models Learned simulators that predict future observations, so a policy can plan or train inside imagination instead of on hardware. |
88 | 7.55% | +2.85pp | |
| Tactile & force sensing Touch as a first-class modality: skins, current sensing, and contact-rich feedback loops. |
78 | 6.69% | +2.2pp | |
| Foundation models & pretraining Cross-embodiment pretraining, scaling laws, and the generalist-policy bet. |
73 | 6.26% | +1.9pp | |
| Simulation, sim-to-real & benchmarks Simulators, domain randomisation, and the evaluation harnesses that decide whether a policy really works. |
70 | 6% | +1.59pp | |
| Humanoids & locomotion Whole-body control, legged locomotion, and loco-manipulation on human-form robots. |
69 | 5.92% | +1.31pp | |
| Data collection & teleoperation The unglamorous bottleneck: rigs, interfaces, and pipelines that produce robot demonstrations. |
79 | 6.78% | +0.91pp | |
| Egocentric & human data Turning first-person human video and wearable capture into robot training signal, bypassing teleoperation. |
53 | 4.55% | +0.85pp | |
| Dexterous manipulation Multi-fingered hands, in-hand reorientation, and contact-rich control of everyday objects. |
26 | 2.23% | +0.75pp | |
| Video & generative modeling The upstream stack robotics borrows: video generation, diffusion and autoregressive models of pixels, and the representations they learn. |
85 | 7.29% | +0.73pp | |
| Safety & evaluation Does it fail safely, and how would we know? Verification, red-teaming, and honest metrics. |
51 | 4.37% | +0.63pp | |
| Hardware & morphology co-design Designing the body and the learning algorithm together instead of in sequence. |
42 | 3.6% | +0.48pp | |
| 3D & spatial reasoning Reconstruction, scene representation, and the geometry a policy needs to act in a room it has never seen. |
38 | 3.26% | -0.16pp | |
| Imitation & diffusion policies Behaviour cloning at scale, and the generative-model action heads that made it work. |
73 | 6.26% | -0.32pp | |
| Human-robot interaction Robots among people: social robots, collaboration, assistance, and the interfaces between human intent and robot action. |
53 | 4.55% | -0.4pp | |
| Navigation & mobility Getting a body from A to B: mobile manipulation, exploration, and outdoor autonomy. |
226 | 19.38% | -0.83pp | |
| Reinforcement learning & control Online and offline RL, model-predictive control, and the safety filters wrapped around them. |
179 | 15.35% | -1.44pp |
Two signals, deliberately kept apart. In text counts papers whose title, abstract or arXiv comment names the organisation, which includes papers about someone else's robot. In author blocks counts the 31 papers we hold full text for and reads the affiliations directly. Neither is an affiliation database, and the second is a small sample.
| Organisation | In text | In author blocks | Total |
|---|---|---|---|
| Unitree | 63 | 3 | 66 |
| NVIDIA | 48 | 7 | 55 |
| Google DeepMind | 18 | 6 | 24 |
| Physical Intelligence | 15 | 2 | 17 |
| Stanford | 4 | 9 | 13 |
| Georgia Tech | 1 | 7 | 8 |
| Meta FAIR | 6 | 2 | 8 |
| UC Berkeley | 1 | 7 | 8 |
| AgiBot | 6 | · | 6 |
| MIT | 3 | 1 | 4 |
| CMU | 1 | 3 | 4 |
| ByteDance | 2 | 1 | 3 |
| Toyota Research Institute | 2 | · | 2 |
| ETH Zurich | · | 2 | 2 |
| Boston Dynamics | 1 | 1 | 2 |
| Columbia | · | 1 | 1 |
| UCSD | · | 1 | 1 |
| Shanghai AI Lab | 1 | · | 1 |
| Amazon | · | 1 | 1 |
| Applied Intuition | 1 | · | 1 |
| Princeton | 1 | · | 1 |
The most common two-word and three-word technical phrases across 2,209 indexed abstracts. This is a description of the corpus on this site, not a measurement of the field. Use the share chart above for that.
One arXiv API query per (topic, month) reading opensearch:totalResults for cat:cs.RO restricted to that month. Share is topic count / all cs.RO submissions that month. Queries are printed next to each topic so anyone can re-run them.
Bigrams and trigrams from titles and abstracts of the papers held in this corpus. Lift compares the term rate in the newest year against all earlier years in the same corpus. Small corpus, so treat as a pointer, not a measurement.
Two independent signals kept separate: (1) organisation names appearing in paper text or arXiv comments, (2) organisation names appearing in the author block of the papers we hold full text for. Neither is a complete affiliation database.
Generated 2026-08-28. Download the raw numbers.