Robotics Papers

2025-07-25 · 56 citations · club pick

Back to the Features: DINO as a Foundation for Video World Models

Federico Baldassarre, Marc Szafraniec, Basile Terver, Vasil Khalidov, Francisco Massa, Yann LeCun, Patrick Labatut, Maximilian Seitzer, Piotr Bojanowski

No peer-reviewed venue on record yet. 56 citations, 5 of them influential, as of the last refresh.

Abstract

We present DINO-world, a powerful generalist video world model trained to predict future frames in the latent space of DINOv2. By leveraging a pre-trained image encoder and training a future predictor on a large-scale uncurated video dataset, DINO-world learns the temporal dynamics of diverse scenes, from driving and indoor scenes to simulated environments. We show that DINO-world outperforms previous models on a variety of video prediction benchmarks, e.g. segmentation and depth forecasting, and demonstrates strong understanding of intuitive physics. Furthermore, we show that it is possible to fine-tune the predictor on observation-action trajectories. The resulting action-conditioned world model can be used for planning by simulating candidate trajectories in latent space.

Ten-minute slide kit

Six slides is the whole talk: what was broken, what people tried, what these authors did, what the numbers say, where it falls over, and the sentence people should remember.

SLIDE 1

The problem

In 2018, Ha and Schmidhuber [\[1\]](#page-9-0) popularized the concept of a *world model*, a neural network that predicts the future state of an environment given past observations and actions taken by an agent. Recently, the subject of world models has gained traction \[2[–9\]](#page-9-2), with conditional generative models showing impressive results on specialized domains such as driving \[2, 7, [9\]](#page-9-2), or video games \[5, [6\]](#page-9-5). Likewise, large-scale generative video models with other kinds…

SLIDE 2

What came before

World models are an active area of research. This section attempts to clarify definitions and organize the heterogeneous landscape of world models. At a high level, we define world models as models capable of predicting the temporal evolution of an environment given past visual observations and an optional conditioning

SLIDE 3

The method

Federico Baldassarre Marc Szafraniec Basile Terver Vasil Khalidov Francisco Massa Yann LeCun Patrick Labatut Maximilian Seitzer Piotr Bojanowski Meta FAIR We wish to train a world model capable of understanding the temporal dynamics of real-world videos. Figure 1 outlines the main components of our method, a *frame encoder* and a *future predictor*. In Section 3.1, we introduce notation for the observation space, *i.e*. video pixels, and we identify the state representation to be modeled, namely patch features in…

SLIDE 4

What they measured

In this section, we empirically validate the quality of our world model. In Section 4.1, we evaluate the latent predictions of an unconditional model on dense feature forecasting tasks and on three physics understanding benchmark. Then, in Section 4.2, we analyze the effect of the different components of our model. Finally, Section 4.3 demonstrates fine-tuning the world model on agent trajectories and applying it to planning on three simulated RL

SLIDE 5

Where it breaks

Not recoverable from the parsed text. Read this section in the paper yourself.

SLIDE 6

One-line takeaway

Keynote 2 by by Arjun Subramaniam (Factory Intelligence)

Assembled from the paper's own PDF, parsed with its layout intact, 100,063 characters of it, then split on the paper's own section headings. Extractive, not generated: every sentence here is lifted from the paper. The takeaway line is the club's own one-liner from its reading list. Check it before you present it.

Figures worth putting on a slide

  • Figure 1: Latent video world model architecture. A frozen DINOv2 *encoder* maps video frames to patch tokens in latent space. The *predictor* is a stack of cross-attention blocks t
  • Figure 2: Autoregressive predictions. For each video, from top to bottom: frames with timestamps, encoder features, autoregressive predictions in latent space. The predictor has ac
  • Figure 3: How far can the model predict? Cityscapes segmentation forecasting performance as context frames are progressively shifted back in time, further away from the target fram
  • Figure 4: Pre-training dataset statistics. For our 66M video dataset, we report the joint histogram of height *vs*. width with highlighted aspect ratios 16:9, 1:1, and 9:16, as wel
  • Figure 5: Unconditional autoregressive rollouts in latent space. For each clip, we feed the model a few initial frames, either 4 or 6, as processed by the encoder. We then roll out
  • Figure 6: Visualization of cross attentions. We visualize the cross-attention of a single query, to all patches of all previous frames, for two intermediate blocks of the predictor

Presented at

Saturday, May 23, 2026
Robotics & World Models Reading Club 09: CVPR Warm-up & Founders Spotlight — DeltaWorld + VisuoTactile Dexterous Hands | San Francisco 0523
Listed on the event page as “A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens”. The arXiv title above is the record.
San Francisco, CA
Why the club picked it. Keynote 2 by by Arjun Subramaniam (Factory Intelligence)

Read next

2025-06-11
Mido Assran, Adrien Bardes, David Fan +27 · 630 citations
2023-04-14
Maxime Oquab, Timothée Darcet, Théo Moutakanni +23 · 10,015 citations
Trans. Mach. Learn. Res.Foundation models & pretraining
2018-03-30
Pauline Luc, Camille Couprie, Yann LeCun +1 · 100 citations
ECCV

Something wrong on this page? Open a correction.