The 3D rasterization pipeline
model→world→view→clip, the division by depth (perspective), near-plane clipping and rasterization, where the z-buffer decides "who is closer" per pixel. This is the skeleton shared by Quake, the GPU and differentiable renderers.z into the w coordinate, and then comes the perspective divide (dividing by w) → normalized coordinates in [−1,1]³, which are scaled into pixels. Before the divide, near-plane clipping is mandatory (you can't divide at z≤0). Depth between pixels is resolved by the z-buffer: for every pixel you store the nearest depth and draw a fragment only if it is closer. Painter and BSP sort polygons; the z-buffer sorts pixels — which is why it displaced almost everything else.
The matrix chain: five spaces
"Where on screen is this vertex?" is a sequence of coordinate system changes, each one a multiplication by a 4×4 matrix in homogeneous coordinates:
Model places the mesh in the world (the instance's position, rotation and scale). View is the camera's inverse matrix: instead of "moving the camera" we move the whole world as if the camera sat at the origin looking along an axis. Projection doesn't project by itself — it prepares the division: it puts the depth z into the w coordinate and scales by the field of view. On a GPU the first three are usually merged into a single MVP matrix (model-view-projection) and applied in one pass in the vertex shader.
Perspective = division by depth
Distant objects are smaller because the screen coordinate is proportional to 1/z. A pinhole camera with focal length d: a point (x, y, z) in camera space lands on the screen plane at
A worked example. Focal length d=4. The point A=(2, 1, 5) → x_s = 4·2/5 = 1.6, y_s = 4·1/5 = 0.8. The same point twice as far away, z=10 → x_s = 4·2/10 = 0.8. Double the depth and you halve the screen offset. That is foreshortening, and it is nonlinear in z — hence all the depth effects (and the precision problems further down).
To pack that division into the matrix chain, you move to homogeneous coordinates: a vertex is (x, y, z, w), and the point (x, y, z, w) "means" (x/w, y/w, z/w). The projection matrix is arranged so that z ends up in the output w (in a right-handed system, −z):
clip = Projection · (x, y, z, 1)ᵀ x_clip = x / (aspect · tan(fov/2)) // horizontal scale y_clip = y / tan(fov/2) // vertical scale z_clip = A·z + B // A,B depend on near/far w_clip = −z // ← this is where depth went NDC = clip / w_clip // perspective divide: dividing by −z → x_ndc, y_ndc, z_ndc ∈ [−1, 1]
Dividing by w is the perspective divide — the only nonlinear step in the entire chain. After it the coordinates lie in the NDC cube [−1,1]³ (Normalized Device Coordinates), and the final viewport transform linearly stretches x_ndc, y_ndc into window pixels and z_ndc into a z-buffer value.
Why clip — and why before the divide
At z=0 the division blows up, and at z<0 (behind the camera) a point with negative w gets mirrored into the frame by the division — geometry behind you "crawls out" in front, upside down. So before the perspective divide, triangles are clipped against the near plane (and, while we're at it, the other frustum faces). A polygon crossing the plane is cut by the Sutherland–Hodgman algorithm: walk the edges, keep the points inside, and where an edge crosses the plane insert a new vertex. A triangle can become a quad or a pentagon — which is then re-triangulated.
The key: clipping is done in clip space, before the divide by w, where the frustum planes are simple linear inequalities −w ≤ x ≤ w, −w ≤ y ≤ w, −w ≤ z ≤ w. After the divide the sign of w is lost, and points behind the camera can no longer be told from points in front. The cheap common case is the guard band: fully visible and fully rejected triangles skip the expensive clipping, and only those crossing the edge get cut.
Rasterization and the z-buffer
After projection a triangle is three points on the screen. The rasterizer walks the pixels inside it (testing the sign of the edge functions / barycentrics) and for each one interpolates depth. Who is closer when they overlap? The answer came from the z-buffer (Catmull and Straßer, 1974): a separate buffer the size of the screen, each cell holding the depth of the nearest fragment so far.
// at the start of a frame: z_buffer[*] = +∞ (nothing farther than that)
for each triangle:
for each pixel (px,py) inside:
z = interpolate_depth(px, py)
if z < z_buffer[px, py]: // closer than what's already there?
z_buffer[px, py] = z // update the depth
frame_buffer[px, py] = color // and write the pixel
// otherwise the fragment is hidden — discard it
A worked example. A wall at z=7 has already filled the pixel (z_buffer=7). A pillar arrives at z=5: 5 < 7 → draw it, z_buffer=5. Then a smoke fragment at z=9: is 9 < 5? no → discard. The order of arrival doesn't matter — the nearest wins. That is the power: no need to sort polygons as in painter / BSP. The price is memory for a full-screen depth buffer plus a read and write per fragment. This is exactly why the PS1 (1994) couldn't afford a z-buffer and why Quake in software got by with span sorting — more on that below.
Backface culling throws away half the triangles before rasterization: after projection, a visible front face has its vertices going around the screen in an agreed winding order (clockwise, say); if the sign of the signed area in screen coordinates is the opposite, the face points away from us and gets discarded. One sign instead of drawing — which is why culling is done in 2D after projection rather than with a normal dot product in 3D.
Deep end · theory: the projection matrix, w=−z and why the z-buffer goes "blind" in the distanceskippable
Why the division is hidden inside a matrix
A 4×4 affine matrix can't divide — it is linear. The homogeneous-coordinate trick: let the matrix's last row be not (0 0 0 1) but (0 0 −1 0). Then the output has w_clip = −z, and the subsequent "normalize w to 1" (the perspective divide) automatically divides x,y,z by depth. Nonlinear perspective thus becomes a linear multiplication plus one shared division at the end — exactly what suits a pipeline and a GPU.
The distribution of depth precision
What gets written into the buffer is not z but a quantity monotonic in 1/z. To map [near, far] → [0,1]:
Check: b(n)=0, b(f)=1. But b depends on 1/z → precision bunches up near the camera. Take n=0.1, f=1000: the level b=0.5 is already reached at z≈0.2. That is, half of all buffer values are spent on the nearest 0.1…0.2, while the whole remainder 0.2…1000 shares the other half. Hence z-fighting — the flicker between two nearly coincident distant surfaces that ran out of bits to be told apart.
The cure: reversed-z
The modern technique: a float depth buffer plus swapping near and far (near → 1.0, far → 0.0). The 1/z hyperbola and float's uneven density near zero cancel into nearly linear relative precision across the whole scene. One of the cheapest rendering upgrades there is: the same data, an inverted comparison, several times less z-fighting.
Deep end · engineering: the order of culling, early-z and fighting z-fightingskippable
- The culling cascade (cheap→expensive): frustum culling of whole objects (by bounding box) → backface culling (the sign of the area) → near-plane clipping (only for triangles crossing the plane) → rasterization. Each layer removes its share before the expensive one kicks in.
- Early-Z: the GPU runs the depth test before the pixel shader if the shader doesn't write depth itself — a hidden fragment never pays for expensive shading. Which is why scenes are often drawn roughly front-to-back or with a separate z-prepass: fill depth first, then shade only the winners.
- Transparency breaks the z-buffer. A semi-transparent fragment has to blend with what is behind it, but the depth test is binary (visible or not). So opaque geometry is drawn with the z-buffer in any order, and transparent geometry in a separate pass sorted back-to-front (this is exactly where the painter's algorithm comes back).
- Z-fighting is cured by spreading near and far apart (don't set near=0.001 "just in case"), by
polygon offsetfor coplanar decals and by a reversed-z float buffer. - Perspective-correct interpolation: attributes (UVs, color) are linear not in screen space but in 3D — you interpolate
attr/wand1/wand divide at the pixel. The same reason Quake interpolatedu/z, 1/z(see the Quake lesson) and why the PS1, lacking it, made textures "melt".
w) puts everyone into screen pixels. The z-buffer is a nearest-first queue at every pixel window: the window remembers who is currently closest and only admits a newcomer if it stands closer still. Nobody sorts anybody in advance — at each pixel the nearest simply wins.
w, that clipping must happen before the divide, and that the z-buffer sorts pixels rather than polygons means reading any graphics bug (inverted geometry, z-fighting, swimming textures, vanishing transparency) as a consequence of a specific pipeline stage rather than as magic.
w).
Systems / data: the MVP chain = composing transformations in an ETL pipeline and fusing them into one pass (just as merging matrices saves multiplications). The z-buffer = "keep-max/keep-min by key" — dedup by latest version, LWW registers in CRDTs, top-1 per partition with no full sort.
ML / AI: the matrix chain is exactly the graph of linear layers that differentiable renderers (PyTorch3D, nvdiffrast) backprop through: the camera projection is a network layer. The projective camera x/z lives on in NeRF and 3D Gaussian Splatting — there the perspective divide is differentiable, and the Gaussian rasterizer is a "soft" z-buffer with alpha compositing. The depth test z<buf itself = max-pooling / scatter-reduce (an argmax over depth = an argmax over a channel). Homogeneous coordinates = the projective embedding in camera calibration, SfM and epipolar geometry.
Performance: early-z = exiting an expensive computation early on a cheap predicate (short-circuiting, predicate pushdown in SQL — drop the row before the heavy join). The guard band = a fast path for "entirely inside / entirely outside" and a slow one only for the boundary case.
The principle: fuse the linear steps into one; defer the nonlinearity to the end; and don't sort everything when each cell only needs to keep the winner.
🕹 Games to play — and what to notice
One pipeline, but every piece of hardware makes its own compromise at the "depth", "vertex precision" and "textures" stages. The era's artifacts are a direct X-ray of which stage got cut. For each: what is inside and what to play to see it with your hands.
The most instructive case — three stages cut at once. (1) No z-buffer: depth was sorted per polygon through an ordering table (the GTE computed the average depth of a triangle's vertices) → at similar depths polygons pop in front of each other. (2) Fixed-point only, no subpixel precision → vertices jitter and jump as the camera moves. (3) Affine textures with no perspective correction → textures swim and wobble on large polygons. The three famous "PS1 glitches" = three pipeline stages thrown out.
🎮 Play: run Crash Bandicoot, Tomb Raider or Final Fantasy VII (an emulator or a PS Classic). Pan the camera along a long wall — you'll see vertex jitter and texture wobble; watch the floor/wall junctions — fragments flicker over which is in front (no z-buffer). That is a pipeline with three holes in it.
The direct contrast to the PS1: the RDP did a per-pixel z-buffer, perspective-correct textures, trilinear mipmapping and antialiasing — all in hardware. Geometry doesn't jitter, textures don't swim. The compromise moved elsewhere: a tiny 4 KB texture cache (TMEM) → small, blurry textures, plus the signature "vaseline" AA blur.
🎮 Play: in Super Mario 64, pan the camera along the same kind of wall — the vertices sit rock still and the texture doesn't wobble (there is a z-buffer and correction), but the texture itself is muddy and stretched across a large polygon (4 KB of cache). Compare side by side with the PS1 and you'll see who kept which stage and who cut it.
A third path: a full-screen z-buffer is expensive on a Pentium, so id Software drew the world (the BSP) by sorting spans by 1/z through an edge list — every world pixel is drawn exactly once, with no depth buffer. But the models (monsters, weapons), sprites and particles did go through a z-buffer, so they intersect the world correctly. A hybrid: the expensive z-buffer only where polygon sorting can't cope.
🎮 Play: in QuakeSpasm type r_speeds 1 — the counter of drawn world surfaces. Thanks to the edge list the world is drawn with almost no overdraw (the counter is modest). Walk right up to a monster standing by a wall — it correctly intersects the geometry: that is the models' z-buffer on top of the world's spans.
Hardware acceleration made the z-buffer cheap → it started covering everything, and the manual tricks with span sorting and FDIV disappeared. What surfaced instead was a purely buffer-side artifact: z-fighting on coplanar surfaces in the distance (not enough depth bits — see the deep end on precision).
🎮 Play: in GLQuake (or any late-90s Voodoo game) find a distant wall with a decal or poster right against it — at range you'll see them flicker against each other as you move. That is a z-buffer that ran out of precision specifically in the distance (the precision went to the camera).
The same pipeline, but visible through dev tools: you can render the depth buffer itself, wireframe (triangles after clipping), an overdraw heat map (how many times a pixel was covered).
🎮 Watch: in Godot, turn on the Overdraw and Wireframe debug draws; in Unity, the Rendering Debugger / Frame Debugger. Pan the camera: you'll see backface culling shaving off the rear faces and the z-buffer killing the overlaps. The pipeline from this lesson, on screen.
u/z, 1/z) and span sorting of the world instead of a z-buffer: the same pipeline, hand-optimized for the Pentium.Why is perspective a division by z rather than a subtraction or something linear?
(x,z) meets the screen plane at distance d at the point x·d/z — a ratio of legs. This is a projective operation, not an affine one: lines parallel in the world converge at a vanishing point precisely because of the division by depth. A linear (affine) projection is orthographic: depth doesn't affect size and nothing converges. The division by z is the mathematical essence of "farther is smaller".Why have a homogeneous coordinate w at all if you could just divide by z by hand?
MVP. If you divide by hand in the middle of the chain, everything before the division can't be fused with the matrices after it, and you can't clip in a convenient linear space. w moves the division to the very end and makes it one shared step for x,y,z. Bonus: w carries the sign of depth — it is what lets you cut away points behind the camera before the division "folds" them into the frame.The z-buffer stores the nearest depth — so why is precision worse for distant objects rather than near ones?
1/z, not z (that falls out of the divide by w). The function 1/z changes steeply near the camera and is nearly flat in the distance → equal buffer steps correspond to microscopic depth intervals up close and enormous ones far away. With n=0.1, f=1000, half the buffer's values are spent on z<0.2. Hence z-fighting specifically on distant surfaces. The cure is a reversed-z float buffer (see the deep end).If the z-buffer is so good, why did Doom and Quake bother with BSPs and span sorting?
The painter's algorithm and the z-buffer do the same thing — why do both still exist?
Backface culling by screen area — what about a single-sided (paper-thin) mesh?
- Scratchapixel — "Rasterization: a Practical Implementation" and "The Perspective and Orthographic Projection Matrix" (derived from scratch).
- Akenine-Möller, Haines, Hoffman, "Real-Time Rendering" — the pipeline, clipping, the z-buffer, reversed-z.
- Michael Abrash, "Graphics Programming Black Book" — Quake's software rasterizer and span sorting.
- Pikuma, "How PlayStation Graphics & Visual Artefacts Work" — why the PS1 jitters and swims (a pipeline missing three stages).
- nvdiffrast / PyTorch3D — differentiable rasterization: the same pipeline, backpropagated through.
- Module 3 (
03-3d-revolution-1993-1999.md), the "3D Graphics Pipeline" section.