Robotics Papers

Topic

Vision-language-action

Policies that consume pixels and language and emit robot actions directly, usually fine-tuned from a VLM backbone.

139papers in 2026-07
11.92%of all cs.RO that month
+7.1ppshare change, last 6 vs prior 12 months
24-month volume

arXiv query: cat:cs.RO AND (abs:"vision-language-action" OR abs:"VLA")

Presented at reading clubs

These already have a slide kit, so they are the fastest ones to pick up and present.

2026-01-28
Brian Y. Tsui, Alan Y. Fang, Tiffany J. Hwu · 2 citations
2025-12-27
Simar Kareer, Karl Pertsch, James Darpinian +5 · 39 citations

Recent work

2026-08-27
Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste +2
2026-08-26
Sanghwan Jang, Minjin Jeon, Minsoo Kim +3 · 1 citation

269 papers in this corpus carry the vision-language-action tag. How tagging works.