World Lab / Interactive talk companion

One present. Many possible futures.

03 / ACTION → CONSEQUENCE

A scene tells you what is here. A dynamics model asks what happens next—if you act. Steer this tiny world, compare counterfactual paths, and watch uncertainty grow with the horizon.

TOP-DOWN WORLD · FIXED TOY DYNAMICS
TIME 000
━ selected action┄ alternative actions◌ perturbed model rollouts
01 / OBSERVE

Observation oₜ

The rendered top-down scene.

pixels → state
02 / REPRESENT

State sₜ

Position and heading, read directly here.

03 / INTERVENE

Action aₜ

Turn rate applied at each step.

04 / PREDICT

Next state ŝₜ₊₁

Apply the same transition once.

A world you can navigate ≠ learned dynamics

A spatial scene can support navigation and novel views without predicting the consequences of actions. In the action-conditioned sense illustrated here, a world model predicts future states or observations from state, history, and actions. Video prediction can model change, but an action input makes interventions explicit.

What this demonstration actually does

This is deterministic, hand-written 2D motion with idealized reflections. There is no trained AI, learned encoder, or physically realistic simulator. The faint paths are a sensitivity ensemble, not calibrated confidence intervals. Near a collision, small parameter changes can cause large differences in the future.

Open the model notebook

At each step: θ′ = θ + 0.035a; x′ = x + v cos θ′; y′ = y + v sin θ′, with v = 3 world units/step and a ∈ {−1, 0, +1}. Walls and circular obstacles reflect the velocity. The state-to-pixel mapping is a renderer, not a learned decoder. Faint trajectories use deterministic speed and turn-rate perturbations controlled by the mismatch slider.

In Ha & Schmidhuber's learned World Models architecture, vision compresses observations to zₜ and a recurrent dynamics model predicts P(zₜ₊₁ | aₜ, zₜ, hₜ). Our direct state sₜ replaces those learned representations only for this illustration.