← Module 11/World models
RU
Module 11 · The modern frontier

World models

A neural network that doesn't render a world from assets but generates it frame by frame in response to your actions. Research frontier, not production.
write-up~15 min
The gist in 20 seconds
A world model is a video model trained to predict the next frame conditioned on the player's action. You feed it key presses and it "continues" an interactive world that exists in no asset file. By 2026 this is done by Genie 3 (DeepMind, Aug 2025: real time ~24 fps, 720p, coherence measured in "minutes"), Muse/WHAM (Microsoft, Feb 2025, trained on Bleeding Edge, published in Nature) and startups like Decart (Oasis → Lucy). It's a "neural engine": impressive, but with no persistent state, no determinism and no precise controllability — a demo for now, not a game.

What it does

A classical engine: assets + physics + code → render a frame. A world model throws all of that out: a single neural forward pass, conditioned on the history and the current action, produces the next frame. The world "lives" in the model's weights. They're trained on huge volumes of gameplay (Muse on ~7 years of Bleeding Edge matches). The key milestones:

ModelWho / whenWhat it showed
Genie 3Google DeepMind, Aug 2025Real-time navigation of a generated 3D world, ~24 fps / 720p, coherence of ~a minute; "Project Genie" opened to AI Ultra subscribers (Feb 2026, research only)
Muse / WHAMMicrosoft, Feb 2025World-and-Human-Action Model (WHAM-1.6B), trained on >7 years of Bleeding Edge matches, ~1 fps — a paper in Nature. Later WHAMM! (Apr 2025, MaskGIT ~500M+250M) — >10 fps at 640×360, playable Quake II in the browser, trained on ~1 week of tester data
Oasis → LucyDecart, 2024–25A real-time interactive generative video model / video transformation

🕹 Games to play — and what to notice

Strictly these are demos, not games — but getting your hands on them is the best way to see both the magic and the ceiling. From something browser-based you can play right now to a research preview you can only watch. The thing to observe is the same everywhere: where the coherence falls apart.

Oasis Decart + Etched · browser, playable

The most accessible one: a diffusion transformer trained on Minecraft video, taking keyboard/mouse input and generating every frame autoregressively — no engine, no assets, no game code. The 500M weights are open (HuggingFace Etched/oasis-500m); the browser demo at ~360p is a cloud stream (generation runs on a server GPU, you don't need top-end hardware), and the weights can also be run locally on a consumer GPU. It understands building, lighting, the inventory — but all of it is "hallucinated per frame".

🎮 Play: open the Oasis demo in a browser, dig out a block, turn away and look back — the world is resampled, the hole is gone. Look up and down a few times and the biome will "drift". Time how many seconds it takes for coherence to break: that's error accumulation in your hands.

Muse / WHAMM Microsoft · browser, Quake II

Microsoft's line in two steps. WHAM (World-and-Human-Action Model) was trained on >7 years of Bleeding Edge matches and published in Nature (Feb 2025) — but it runs at ~1 fps. Then WHAMM! (Apr 2025) swapped autoregression for MaskGIT (generate the tokens of the whole frame at once, then mask and refine) → >10 fps at 640×360 and playable Quake II in the browser, trained on only ~1 week of tester data on a single level. The same class, a different donor engine and a different decoding architecture.

🎮 Play: run the WHAMM! demo, shoot and move around — notice the low fps and the "liquid" walls that change geometry when you come back. Compare the feel with real Quake II: that's the price of a "neural engine" against rasterization.

Genie 3 DeepMind · research preview, watch only

The most powerful and the least accessible: real-time navigation of a 3D world generated from a text prompt, ~24 fps / 720p, coherence of "minutes" (announced Aug 2025). "Project Genie" was cracked open to AI Ultra subscribers (Feb 2026), but it's research only, not a product.

🎮 Watch: go through the official Genie 3 clips frame by frame — look for the moment where an object leaves the frame and comes back different (no persistent state), and estimate how many seconds the coherence holds. That's the 2026 high-water mark — and its limit.

Deep end: autoregression, error accumulation and why coherence falls apartskippable

The model parameterizes

p(x_t | x_<t, a_<t)

and generates by rollout: it feeds its own outputs back into the input.

Exposure bias / error accumulation

It was trained on the true history (teacher forcing), but at inference it conditions on its own (imperfect) frames → distribution shift. The per-step error ε isn't damped, it accumulates: in the naive imitation limit the divergence grows as ~O(ε·T²) over horizon T (the DAgger analysis), versus O(ε·T) with on-policy correction. That's the mathematical reason coherence holds for "minutes" and not hours.

Two different meanings of "world model"

  • Generative video (Genie, Muse): p(frame | history, action) in pixel space — the goal is realism/playability.
  • Model-based RL (Ha & Schmidhuber 2018, "World Models"; Dreamer): a latent dynamics model for planning/imagination — the goal is usefulness for control, not a pretty frame.

Why persistent state isn't "just add memory"

A pure autoregressive video model has no explicit state store: its "memory" lives in a finite context window. Facts outside the window aren't represented → come back to a place and it gets resampled from scratch. Persistence requires an explicit state structure (a hybrid with a simulation/database), not just a longer context.

Deep end · economics and infrastructure: why this isn't shippable yetskippable

Beyond the incoherence there are prosaic blockers:

  • Cost per frame: neural frame generation is orders of magnitude more expensive than rasterization — running it 60 times a second per player is economically absurd today (a datacenter GPU per session).
  • Latency and infrastructure: real-time inference needs expensive hardware close to the player; on a consumer GPU you compromise on quality/fps.
  • The design question: which genre actually wants a non-persistent, non-authored world? For now it's a prototyping/research tool and "neural modding", not a product — the business case isn't closed.
Analogy
A normal engine is a puppet theatre: the sets, the puppets and the strings are made in advance, and the code pulls them. A world model is an artist who paints the next frame of a dream in response to your action. While you walk forward, the corridor gets "invented" on the fly. Turn away and come back and the wall is different: a dream has no persistent memory. That is both the magic and the central flaw.
Why it matters
This is a potential paradigm shift in rendering — a world with no asset pipeline. But for an engineer in 2026 the value is a sober assessment: it's a research breakthrough and a tool for prototyping/"neural modding", not a way to make your game. Knowing what Genie 3 / Muse can do, and what they still can't, is part of a professional picture of the frontier.

What doesn't exist yet

🔁 Beyond games — where this transfers
The main mechanism is autoregression and error accumulation (exposure bias), plus the distinction "a generative model ≠ a simulator with state". It carries straight over into ML:

ML / AI: the same error accumulation in any autoregressive rollout — long LLM generations "drift"; teacher forcing vs free running; why DAgger fixes O(εT²)→O(εT). World models in model-based RL (Dreamer) = an "imagined" simulator for planning.

Robotics / control: sim-to-real and compounding error in dynamics prediction — the same horizon problem.

Principle: per-step error compounds over the horizon; "looks plausible" ≠ "has persistent state" — a general limit of generative models.

🔧 Run it and poke at it — on your home machine
What to play/watch is above (🕹). Here it's about getting into the model itself:
🔧 Poke at it (debug) ~1 h, Python + GPU
Get open-oasis running locally (repo etched-ai/open-oasis + weights Etched/oasis-500m, a GPU is required). Feed it your own action sequence and look at the rollout frame by frame. Shrink the context window in the inference script — the coherence breaks sooner (there's the dependence of horizon on window). Freeze the input on a single frame and you'll see the model keep "painting" drift anyway.
🧪 Test it (with QA eyes) ~15 min
Measure the coherence horizon: seconds/frames until the world falls apart in Oasis or WHAMM!. Run the "turn away and come back" test — is the state fixed (it isn't). Compare the fps with real Minecraft/Quake II — feel the price of generating a frame with a network versus rasterizing it.
Checklist: ran open-oasis locally; shortened the context and saw an earlier collapse; measured the coherence horizon in seconds; confirmed "turn away → the world is different".
Connections
foundation
LLM NPCs — the same "ML generates content" line; an LLM generates dialogue, a world model generates the pixels of the world.
contrast
Classical vs ML — a world model is the extreme point of "ML for content": even the frame itself is generated rather than drawn by an engine.
Questions worth asking
Why does coherence fall apart specifically "after minutes" — mathematically?
An autoregressive rollout feeds the model its own outputs: the per-step error isn't damped, it accumulates (exposure bias / distribution shift). In the naive limit the divergence grows roughly quadratically with the horizon (the DAgger analysis), plus facts outside the context window are lost. Hence "minutes of coherence" rather than hours — see the deep end above.
How does Genie's world model differ from the world model in Dreamer/Ha-Schmidhuber?
Different goals. Genie/Muse are generative video, p(frame|history,action), optimizing realism/playability. Dreamer and "World Models" (2018) are latent dynamics models for planning (the agent "imagines" trajectories and learns inside them). The first is for watching/playing, the second is for control. Same term, different problems.
What's needed for persistent state, and why isn't it "just add memory"?
A pure video model holds "memory" only inside its context window — beyond it the world gets resampled. Persistence requires an explicit representation of state (a structure/simulation/database) that the model is obliged to respect — that is, a hybrid with a classical engine. That's a different architecture, not "a longer context".
How is this different from Sora and video generation at all?
The key is conditioning on action in real time: the model predicts the next frame in response to your input right now, rather than rendering a finished clip. It's exactly that interactive conditioning that turns a video model into a (quasi-)playable environment.
Further reading / watching