← Module 3/Client-side prediction
RU
Module 3 · The 3D revolution (1993–1999)

Client-side prediction: how to hide the ping

QuakeWorld (1996) made online shooters playable over dial-up with one idea: the client draws the result of an input immediately, without waiting for the server, and the server confirms later and corrects any divergence. Authority stays with the server, responsiveness with the client.
deep~16 min
The gist in 30 seconds
If the server is authoritative and the client waits for a reply to every input, movement lags by a full ping (RTT) — on 1996 dial-up that is 200–400 ms, unplayable. Client-side prediction (Carmack, QuakeWorld, August 1996): the client applies the input locally immediately with the same physics as the server, and sends the input with a sequence number. The server is the authority: it simulates and sends back the state plus the number of the last input it processed. The client compares its prediction at that number against the truth; if they match (the usual case) it does nothing, and if they diverge it rolls back to the server state and replays all inputs not yet acknowledged (server reconciliation). Other players can't be predicted (you don't have their input) → they are interpolated from past snapshots, drawn slightly in the past. And so that a shot at a moving target registers fairly, the server rewinds time (lag compensation). Server authority catches cheats; prediction is only presentation.

The problem: authority costs a ping

The naive "the server decides everything" scheme is safe but slow. Press W → the packet flies to the server → the server moves you → the state flies back → only now does the character take a step. Movement latency = a full round trip:

tmove= RTT= tup+tdown

At 120 ms RTT every step, jump and turn is 120 ms behind — controls by post. You can't simply hand power to the client ("I'm at X, I have 999 HP") — that opens the door to cheating. You need both server authority and instant response. The resolution is prediction.

Prediction + reconciliation

The client does two things at once for every input: (1) applies it locally and immediately — the character moves in that very frame; (2) sends the input to the server with a monotonic sequence number. The server processes the input, moves authoritatively, and in the reply snapshot reports "last processed input = N" (the ack) plus the resulting state.

The client keeps a history of its own inputs. On receiving an ack for N it: discards everything ≤ N from the history; takes the server's state as truth; and replays the not-yet-acknowledged inputs N+1, N+2, … on top of it with the same formulas. If the prediction matched the server (which it does in 99% of frames on a decent connection) the position doesn't twitch. If it diverged (the client thought it was running while the server knew about a wall) the position slides gently toward the truth.

Client Server t0: input #50I move at once input+#50 → simulate, ack=50 ← state+ack reconcilematched → 0 corrections the RTT elapsed, but the player was moving from t0 — the ping is hidden

A worked example. RTT 120 ms, the client is on input #50. Without prediction the character would stand frozen until t=120 ms. With prediction it stepped at t=0; at t=120 ms an ack for #47 arrives with the server's position. The client sets its position to the server's (as of #47) and instantly replays #48, #49, #50 → ending up exactly where it was already drawing. The player noticed nothing. If, however, the server saw a wall between #47 and #50, the replay runs into it → a gentle nudge toward the truth instead of a teleport.

A hard requirement — determinism: client and server must run identical movement code. Let the formulas differ by a hair and the prediction will miss constantly, and the character will be jolted by corrections every frame. Which is why player physics is written once and compiled into both sides.

Other players: interpolation in the past

You can predict your own input — it is in your hands. You can't predict anyone else's: you don't know which way the enemy will press. So the others are interpolated rather than predicted between the two most recent snapshots received — and deliberately drawn slightly in the past (an interpolation buffer Δ, typically 50–100 ms), so there are always two points for a smooth lerp:

prender= (1−α)·S(t0) + α·S(t1) , α= t−t0t1−t0

where t = t_now − Δ and S(t0), S(t1) are the snapshots before and after. The alternative is extrapolation (dead reckoning): continue the enemy's motion at their last velocity. Cheap, but if they turned you get a rubber-band snap on correction. Most fast shooters choose interpolation (paying a fixed delay for smoothness) and keep extrapolation for the occasional dropped packet.

Lag compensation: the server rewinds time

Now the conflict: you shoot at an enemy you see in the past (interpolation plus ping). On the server's "now" the enemy has already moved. If the server checks the hit against its own "now", you will miss what you aimed at. Yahn Bernier's solution (Valve, Counter-Strike; the paper "Latency Compensating Methods", 2001): the server rewinds the world to the moment the shooter saw and checks the ray there:

trewind= tnow− (tlat+Δ)

The server keeps a ring buffer of recent positions for every player; for a client's shot it restores the hitboxes at t_rewind (their latency plus their interpolation buffer) and traces there. The price of the compromise is the famous "I was already around the corner but died anyway": from the shooter's point of view you were still in the open, and the server sided with the shooter (favor-the-shooter). Making it fair for both at nonzero ping is mathematically impossible — you choose who gets the truth.

Deep end · engineering: tick rate, snapshots, UDP and packet lossskippable
  • UDP, not TCP. TCP guarantees order and delivery → a lost packet stalls everything behind it (head-of-line blocking), and in a shooter a stale packet is no longer needed — you want the freshest one. So you send UDP and decide yourself what needs reliability (events: "the door opened") and what can be dropped (positions — the next snapshot is coming).
  • Tick rate. The server simulates in fixed ticks (Quake: tens per second; CS:GO: 64/128; Overwatch pushed it to ~63). A higher tick = a more accurate simulation and lag comp, but more CPU and more traffic. The client renders at its own frame rate, with interpolation and prediction filling the gaps between ticks.
  • Snapshots and delta compression. The server sends world state N times per second; to fit the pipe it sends deltas against the last snapshot the client acknowledged (only what changed) plus a priority for nearby and visible entities (the PVS from the Quake lesson cuts network volume too, not just drawing).
  • Input redundancy. Over UDP the client re-sends the last few inputs in every packet, not just the newest one — a single lost packet doesn't punch a hole in the history, and the server picks up the duplicate from the next one.
  • The jitter buffer. Packets arrive unevenly; a small buffer smooths the spread of arrival times at the cost of a little more latency. It is the same interpolation buffer Δ.
Deep end · hosting: authority, determinism and where to compute itskippable
  • The authority model. A dedicated server as the single source of truth is the standard for competitive games (anti-cheat, no "host advantage"). P2P or a listen server is cheaper (no infrastructure), but the hosting player has zero ping and is harder to protect against cheating.
  • Determinism across platforms. If prediction or lockstep relies on the simulation matching exactly, the float problem surfaces: the same code on different CPUs or compilers can give slightly different results (operation order, FMA, x87 vs SSE). Lockstep RTS games pin the math down (fixed-point or a strictly specified float) — otherwise the clients drift apart (desync).
  • Consistency vs latency. This is the same dilemma as in distributed systems: strong consistency (wait for the authority) is safe and slow; an optimistic local action is fast but requires reconciliation and rollback. Game development chose optimism plus reconciliation long before it became mainstream on the web.
  • Regional servers. You can't cheat the physics of ping: RTT ≥ 2·distance/c. So matchmaking ties you to the nearest data center — to lower the base RTT that all of the prediction runs on.
Analogy
Client-side prediction is an optimistic UI: you hit "send" and the message appears immediately (prediction), though the server hasn't confirmed it; if an error comes back the display is rolled back (reconciliation). Or autocomplete: typing continues on a guess and a correction arrives if the guess was wrong. And lag compensation is like a referee reviewing the tape: to award a goal fairly, you look at the frame when the shot was taken, not at "now".
Why it matters
This is the resolution of a hard conflict between authority and responsiveness, which shows up everywhere there is a network and a user expecting an instant reaction. The chain "act optimistically in local, reconcile with the authority, roll back on divergence" is a canonical pattern, and games proved it in the field a decade before optimistic updates became normal on the web and speculative execution became normal in processors and LLMs.
🔁 Beyond games — where this transfers
Prediction plus reconciliation is speculative execution (act on a guess, roll back on a miss) and optimistic concurrency (don't block waiting, verify at the end).

Systems / frontend: optimistic UI updates (React/Redux: show the result before the server answers, roll back on an error) are literally prediction+reconciliation. Optimistic concurrency control in databases (versions/CAS instead of locks). Eventual consistency and CRDTs: the local edit lands immediately, convergence comes later.

ML / AI: speculative decoding is an exact copy of the technique: a cheap draft model "predicts" several tokens ahead, the large model verifies them in one pass and rolls back from the first divergence, replaying onward — that is reconciliation by input number, word for word. Off-policy RL: distributed actors (IMPALA) act on a stale policy while corrective importance sampling (V-trace) fixes the training mismatch — the same "act optimistically, then correct against the authority".

Hardware: branch prediction plus speculative execution in a CPU: the processor guesses a branch and computes ahead, flushing the pipeline on a miss (rollback) — prediction with rollback in silicon.

The principle: don't pay the latency of waiting for the truth — act on a deterministic guess, keep a history, reconcile against the authority and replay only what diverged.

🔧 Run it and poke at it — on your home machine
What to play is below (🕹). Here — see prediction through counters and the console:
🔧 Poke at it (debug) ~30 min, Source/Quake
In a Source game (CS:GO/2, TF2) open the console: net_graph 3 gives you ping, loss, tick rate, interpolation. Play with cl_interp / cl_interp_ratio (the interpolation buffer for other players) and cl_predict 0/1 — with prediction off you'll feel movement arriving "by ping". Turn on sv_showhitboxes 1 and (on your own server) the lag-comp visualization: you'll see the server placing hitboxes where the target was in the past. In a QuakeWorld source port, cl_predict and the net graphs work the same way.
🧪 Test it (QA eyes) ~15 min
Provoke the artifacts: artificially raise ping and loss (net_fakelag, net_fakeloss in Source) and catch rubber-banding when you run into a wall, other players warping under loss, "death around the corner" (lag comp favoring the shooter), desync on a jittery connection. The network artifact checklist.
Checklist: turned cl_predict off and felt the ping; played with cl_interp and saw smoothness versus delay; reproduced "death around the corner" and understood why the server sides with the shooter. A deep [LAB] simulator of lag and prediction is in Module 4.

🕹 Games to play — and what to notice

From the first implementation in QuakeWorld to the industry standard and the contrasts (lockstep, rollback). For each: what is inside and what to play / what to type into the console.

QuakeWorld 1996 · where it all started

The first widespread implementation of client-side prediction (Carmack, .plan of 16 Aug 1996: letting the client guess the outcome of movement until the server's authoritative answer arrives). It made Quake playable over dial-up and gave birth to the competitive online shooter.

🎮 Play: in a source port (ezQuake/FTE QuakeWorld) set a high ping and toggle cl_predict 0 vs 1 — at zero, movement lags by the ping; at one it is instant. You are flipping exactly the switch Carmack added in '96.

Counter-Strike lag compensation

Bernier at Valve added lag compensation on top of prediction: the server rewinds hitboxes into the shooter's past. Hence the signature "got around the corner and died anyway".

🎮 Play: in CS, type net_graph 3 and look at ping/tick/interp. Catch the moment you die already behind cover — that isn't a bug, it is favor-the-shooter: on the enemy's screen you were still in the open, and the server rewound to their frame.

Overwatch a high tick rate + favor-the-shooter

Covered in detail at GDC 2017 (Tim Ford): a high tick rate, prediction, interpolation and a deliberate choice to "trust the shooter". A good modern reference for the same architecture.

🎮 Watch: turn on the network overlay in the settings (ping, tick rate). Compare the feel of hitscan on a hero with an instant shot (lag comp decides everything) against a projectile (it travels — partly predicted in flight). The same netcode model, different weapon types.

Fighting games (GGPO) contrast · rollback

The other branch of prediction: rollback netcode (GGPO; Killer Instinct, Guilty Gear Strive). It predicts the opponent's input (usually "they're pressing the same thing as last frame"), simulates onward, and when the real input arrives does a rollback and re-simulation of those frames. No authoritative server — deterministic P2P lockstep with rollback.

🎮 Play: in GGST or any rollback fighting game online, catch the micro-"shudder" of a character during a ping spike — that is visible rollback plus re-simulation. A deep dive on rollback and lockstep is in Module 4.

RTS lockstep contrast · StarCraft/AoE

A completely different choice: with hundreds of units, sending their positions is unrealistic. You send only the commands and run the simulation deterministically and in sync on every machine (lockstep). Latency is hidden not by prediction but by a small command delay (turn delay).

🎮 Play: in old StarCraft/AoE on a bad connection, notice that the whole world lags at once (waiting for every player) rather than one unit rubber-banding — that is lockstep, the opposite of prediction. Why it works that way is determinism and volume (see the deep end).

Connections
foundation
Quake: real 3D — QuakeWorld grew out of the same engine; the PVS cuts not only drawing but network volume too (who gets sent which entities).
next
Cameras — the module's next lesson. And the taxonomy of netcode (client-server / rollback / lockstep) plus a [LAB] simulator of lag and prediction wait in Module 4: the foundation is here, the depth is there.
contrast
Doom (deathmatch) — early networked Doom was peer-to-peer lockstep over a modem: everyone's input to everyone, with no authority and no prediction. QuakeWorld dropped that in favor of client-server.
Questions worth asking
If the server is the authority anyway, why predict on the client at all?
Authority and presentation are different layers. The server decides what happened "for real" (anti-cheat, conflict resolution). Prediction decides what the player sees right now, so the controls don't lag by a ping. Without prediction the game is fair but feels like it arrives by post. With prediction it is responsive, and the server still has the last word through reconciliation. It isn't either/or but two independent questions: "who is right" and "what to show immediately".
Why UDP rather than reliable TCP — packets do get lost, after all?
That is exactly why. On loss, TCP blocks everything after it until the missing piece is resent (head-of-line blocking) — and in a shooter a stale position is already useless, you want the freshest. UDP lets you send without guarantees and decide yourself: positions can be dropped (the next snapshot is coming), while important events (a death, a door opening) get duplicated or acknowledged manually. Reliability applied where it is needed instead of universally and slowly.
Why draw other players in the past rather than extrapolate forward?
To have two real points to interpolate between instead of guessing. Extrapolation (continuing at the last velocity) looks fine until the enemy changes direction — then you get a rubber-band snap on correction, worst of all on sharply maneuvering targets. Interpolation with a Δ buffer pays a fixed delay (50–100 ms), but movement is always smooth and matches what the server actually sent. For hitscan, that delay is compensated by lag comp on the server.
What breaks if the client and server physics aren't identical?
The prediction will systematically diverge from the server → reconciliation will jerk the position every frame (constant jitter or rubber-banding), even on a perfect connection. Which is why player movement is written as one piece of code for both sides, and lockstep games additionally fight float non-determinism across platforms (operation order, FMA) — there a divergence means a full desync and the worlds drift apart. Determinism isn't a luxury but a precondition for prediction working at all.
"I died after I'd already gone around the corner" — is that lag, a bug or design?
A lag-compensation design compromise. The server rewound the hitboxes to the moment of the shooter's shot; on their screen (given their ping and interpolation) you were still leaning out. The server sided with the shooter (favor-the-shooter), because otherwise everyone who sees a target in the past would miss. Being fair to both the shooter and the one running away at nonzero ping is impossible — the developer deliberately decides who gets the disputed frame.
Doesn't prediction open the door to cheating — the client is moving itself, after all?
No, because prediction gives the client no power. The client sends input, not "my position as truth"; the server simulates itself and can reject the impossible (a teleport, speed above the limit, shooting through a wall). The client merely draws in advance what it expects from the server. Cheats are caught by server-side checks and by the fact that the server always recomputes the outcome; prediction only affects what is visible before confirmation.
Further reading