← Module 11/Classical vs ML
RU
Module 11 · Synthesis

Classical vs ML: how to choose

The connection map for the whole module. One sentence worth taking away: the classics are for control, ML is for content.
write-up~15 min
The gist in 20 seconds
Classical AI is for control (FSM, BT, A*, GOAP): 60 FPS, determinism, debuggability, the same NPC today and tomorrow. ML / generative AI is for content (LLM dialogue, diffusion assets, world models): it generates what's too expensive or too static to author by hand. They don't replace each other — they coexist. The question to ask yourself: do I need to control reliably or to produce cheaply?

The decision tree

Control or content generation? control content Map/rules known? Can rules constrain it? yes no A* / FSM /BT / GOAP RL / search(expensive) yes no WFC / BSP /grammar LLM /diffusion control → teal (classical, deterministic) content → pink (generation, costlier the fewer the constraints) Rule: take the dumbest tool that solves the problem.

The module's five patterns

A summary of what recurs across every topic:

🕹 Games to play — and what to notice

One question — "where does strong agent behavior come from" — and the whole spectrum of answers, from zero ML to pure deep RL. The cases are ordered by how much learning is in the agent versus script. Notice where ML actually reaches production and where it stays a research demo.

F.E.A.R. 2005 · zero ML, a pure planner

The benchmark for "smart" AI — and there isn't a single neural network in it. It's GOAP (Goal-Oriented Action Planning, Jeff Orkin): about a dozen actions with pre/post-conditions + an A* planner over the state space. Flanking, suppressing fire, vaulting cover, flipping tables — all emergent from the planner and the loud callouts ("Flanking!"), neither scripted nor learned. Classics-for-control in its purest form.

🎮 Play: fire up F.E.A.R. and fight the Replica soldiers. Notice how they flank and lay down suppressing fire while one of them moves up — and that this is repeatable and debuggable. "Looks like ML" ≠ "is ML".

AlphaGo / AlphaZero 2016–17 · hybrid: search + a learned evaluation

Not "a neural network plays Go" but an MCTS search that a network advises on where to look (policy) and who's winning (value). Remove the search and you drop to strong-amateur level; remove the network and MCTS floods a meaningless tree. AlphaGo beat Lee Sedol 4:1 (March 2016); AlphaZero (2017) reached superhuman play from scratch through self-play, with no human games. This is the canonical "ML makes classical search smarter without replacing it".

🎮 Poke at it: install KataGo (open source, the same architecture) with analysis on — you'll see the MCTS visit tree and the network's evaluation on top. More in the 🔧 below.

OpenAI Five · AlphaStar 2018–19 · pure deep RL at the frontier

No search at runtime — a learned policy (a neural network) emits actions directly. OpenAI Five beat the world champions OG at Dota 2 (April 2019); AlphaStar reached grandmaster in StarCraft II (above 99.8% of players, Nature, Oct 2019). Possible — but the price is wild: OpenAI Five was 256 GPUs + 128,000 CPU cores through PPO, ~250 years of simulated Dota per day, tens of thousands of years in total. Superhuman, and unshippable as a rank-and-file NPC.

🎮 Watch: recordings of OpenAI Five vs OG and AlphaStar replays. Notice the inhuman coordination and micro — and that this is a research demo, not an opponent in a boxed game.

GT Sophy 2023 · RL control that actually shipped

The rarest case of ML control reaching a commercial game. Sony AI trained an RL agent (QR-SAC) to race superhumanly and yet "cleanly", without dirty nudges. First the cover of Nature (Feb 2022) for overtaking champions in Gran Turismo Sport, then a debut for players — the "Race Together" event in Gran Turismo 7, update 1.29 (launched Feb 21, 2023). The conditions under which RL control ships at all: a narrow, precisely specified domain with a rich simulator — and even then it took Sony AI and a Nature paper.

🎮 Play: in GT7, race against Sophy. Notice the clean overtakes and the defensive lines through corners — and that everything around it (menus, events, physics, the rest of the AI) stayed classical. The exception that proves the rule from the tree above.

Deep end: game theory — when it actually applies (and why game AI isn't Nash-optimal)skippable

Under the hood of Utility AI — decision theory

The von Neumann–Morgenstern axioms: if preferences are complete, transitive, continuous and satisfy independence, they can be represented as maximization of expected utility. Utility AI = an approximation of argmax_a E[U(a)] with a hand-written U. That's its formal justification, rather than "scoring by feel".

Where game theory proper applies (several strategic agents)

  • Perfect information, zero-sum, sequential (chess, Go): minimax; the value exists (Zermelo). Solved with minimax+αβ or MCTS+a value net (AlphaGo).
  • Imperfect information (poker): Nash via Counterfactual Regret Minimization (CFR) — regret matching converges to ε-Nash in self-play (Libratus/Pluribus).
  • Simultaneous moves / general-sum (RTS, social games): there can be many Nash equilibria, and computing them in general is PPAD-complete (Daskalakis et al.) — practically intractable.

The complexity of "solving a game"

Generalized versions of board games are usually PSPACE- or EXPTIME-complete (generalized Go is EXPTIME-complete, Robson 1983; generalized geography is PSPACE-complete). So an exact solution is out of reach → heuristic search + a learned evaluation instead of "solving".

Why shipped AI is NOT Nash-optimal

The goal is the player's enjoyment, not victory. Optimal AI is often not fun: too strong, exploitative, illegible. Designers deliberately weaken the AI (rubber-banding, telegraphed attacks). Game-theoretic optimality is needed only where competition is an end in itself: fighting games (frame data), poker bots, rating-based matchmaking (Elo/TrueSkill — a Bayesian skill model, not game theory).

Deep end · economics and engineering: AI as a cost line, not as magicskippable
  • Build vs buy: your own ML feature = a team of ML engineers, data, infrastructure; an off-the-shelf one (Inworld and the like) = a subscription plus vendor lock-in. Is there anyone to maintain it?
  • ML tech debt: the model drifts, the vendor changes the API, you need retraining/monitoring. A classical FSM doesn't rot on its own — ML does.
  • QA of non-determinism: "smart" AI that can't be tested and balanced predictably often costs more than it brings in. Debugging cost is a real line item.
  • Conclusion: ML is justified when generalization/content delivers value you can't produce by hand and the team can carry the operations. Otherwise the cheapest tool that solves the problem wins (see the decision tree above).
Analogy
It's like choosing between a clockwork mechanism and a painter. The clock (the classics) does the same thing precisely, predictably and cheaply — perfect when you need reliable behavior. The painter (ML) draws endlessly varied things, but expensively, slowly and unpredictably — you need one when authoring everything by hand is impossible. The engineer doesn't pick "which is cooler", they put each tool where it's strong.
Why it matters
The main mistake of 2024–26 is dragging LLMs/RL into places the classics do better, because it's fashionable. Maturity is the opposite: take the simplest tool that solves the problem, and reach for ML only where the classics fundamentally can't cope (open-ended dialogue, asset generation, content per player). This page is the filter to run every "let's bolt on a neural network" through.
🔁 Beyond games — where this transfers
The thesis itself — "take the simplest tool that solves the problem; ML is for generalization, not the default" — is general engineering judgment, not a games thing:

ML / AI (your senior skill): a classical baseline before deep learning — logistic regression/gradient boosting often beat a network on tabular data; regex before NER, a heuristic before a model, rules before RL. Knowing when NOT to use ML is what you're valued for at your level at Artificial Agency.

All of engineering: build vs buy, simple vs smart; "don't reach for a neural network / distributed system / microservices while a monolith or a heuristic still carries it".

Principle: optimal ≠ fashionable; taste in choosing tools (and the nerve to choose the boring one) is the main engineering skill.

🔧 Run it and poke at it — on your home machine
What to play is above (🕹). Here it's about touching "search + a learned evaluation" with your hands and running your own case through the decision tree:
🔧 Poke at it (debug) ~40 min, KataGo
Install KataGo (open source, AlphaZero architecture) with an analysis GUI (KaTrain / Lizzie). Turn on MCTS visits and the win rate. Compare max visits = 1 (the pure policy network, no search) against high visit counts (full MCTS) on the same position — feel how much the search itself adds on top of the network. That's "ML makes search smarter" live. Break it: give it a tactical trap position and you'll see the network's snap judgment and the search diverge.
🧪 Run it through the filter (decide) ~15 min
Take three behaviors from games you have installed: a shooter enemy, a racing opponent, an NPC line of dialogue. Run each through the decision tree from the top: control or content? are the rules known? can rules constrain it? Predict "classical / ML" — then check what the studio actually did. That is the filter against "let's bolt on a neural network".
Checklist: saw the policy-only vs MCTS difference in KataGo; ran ≥3 cases through the tree; caught at least one where "fashionable ≠ optimal".
Connections — the module map
control
FSM and Behavior Trees — the core of "classics for control".
control
Pathfinding A*/NavMesh — a known map → the classics beat RL outright.
content+rules
Wave Function Collapse — generation squeezed by constraints: the midpoint of the spectrum.
content
LLM NPCs — "ML for content" in its pure form; the control around it stays classical.
content
World models — the extreme point of generation: the world's frame itself is generated.
Questions worth asking
Why isn't shipped game AI made Nash-optimal?
The goal is fun, not victory. An optimal opponent is usually not fun (illegible, exploitative, too strong), and computing Nash in the general case is also PPAD-complete. So AI is deliberately weakened and made legible (telegraphed attacks, rubber-banding). Game-theoretic optimality is needed only where competition is an end in itself (fighting games, poker bots).
Which algorithm when: minimax, MCTS or CFR?
By the structure of the information and the sum. Perfect information + zero-sum + sequential → minimax/αβ or MCTS (+a value net for large trees, as in AlphaGo). Imperfect information (poker) → CFR (it converges to Nash in self-play). Simultaneous moves/general-sum → Nash, but it can be multiple and PPAD-complete. Classify the game first, pick the tool second.
"Planning ≠ learning" — so what is MCTS in AlphaGo?
A hybrid: MCTS is planning inside a known model (the rules of Go), the value/policy network is learning that makes the search smarter. AlphaGo is about the pairing: learn a heuristic, search with it. The same logic in games: ML strengthens classical search rather than replacing it.
Utility AI — what's its rigorous justification?
The von Neumann–Morgenstern theorem: "rational" preferences can be represented as maximization of expected utility. Utility AI is exactly an approximation of argmax_a E[U(a)] with a hand-written U — so the scoring isn't "by feel", it's an approximation of an EU maximizer (with all the caveats about how U was specified by hand).
Are Elo/TrueSkill in matchmaking game theory?
No, they're a Bayesian skill model: a latent rating + a probabilistic model of the outcome (TrueSkill is a factor graph with Bayesian updates). Game theory is about strategies inside a game, while a rating is about estimating player strength in order to match equals. A common confusion.
Further reading