The render pipeline and culling
The mechanism: the pipeline and its logistics
The pipeline stages
The GPU pushes geometry through a pipeline (simplified):
The GPU is a firehose with high latency
The GPU is not fast on a single thread — it is massively parallel: thousands of threads grouped into warps of 32, executing in lockstep. Memory access is slow (hundreds of cycles), but the latency is hidden by occupancy — while one warp waits on memory, the GPU switches to another. Two killers of throughput: divergence (threads in a warp take different branches → the warp executes both) and scattered access (neighboring threads read non-neighboring addresses → many transactions instead of coalescing). The takeaway: the pipeline wants coherent, batched, predictable work — and all optimization is about that.
Culling — don't process the invisible
The biggest lever is not sending into the pipeline what nobody will see, and doing it as early as possible:
- Frustum culling: an object outside the camera's view frustum → skipped entirely (on the CPU, via bounding volumes).
- Backface culling: a triangle turned away from you (the sign of its screen-space area) → about half of all triangles never get rasterized.
- Occlusion culling: an object hidden behind other geometry → skipped (hi-Z, occlusion queries; the ancestor is PVS in Quake).
Each level cuts work before the expensive stages. In an open world that is the difference between "a million triangles" and "100 thousand visible ones".
Overdraw — don't paint a pixel twice
If you draw far-then-near, the near geometry overwrites the far, and all the fragment-shader work on the far pixel is thrown away. That is overdraw, the quiet performance killer (especially transparency, particles and smoke — they cannot be early-z rejected). The cures are front-to-back sorting (the early depth test rejects occluded fragments before shading) and a depth prepass (render depth only first, then shade only what is really visible). A scene with 3× overdraw shades three times more pixels than it shows.
Draw calls — don't ship one part per truck
Every draw call is a CPU→GPU command with overhead (state changes, driver validation). That is a CPU cost, not a GPU one. The 60 fps frame budget is 16.67 ms; the CPU cost:
10,000 objects at one call each × ~0.05 ms = 500 ms — the frame is impossible (CPU-bound). The cures are instancing (one call draws one mesh N times — grass, crowds) and batching (merge meshes sharing a material): 100 calls × 0.05 = 5 ms — that fits. And explicit APIs (Vulkan/DX12) removed the driver as the single bottleneck: N threads build N command buffers and submit at the end of the frame (in GL/D3D11 the driver was a single-threaded chokepoint).
Scaling lights: forward vs deferred
Naive forward lights every object with every source → (100 lights × 1000 objects = 100k passes). Deferred first writes geometry (normal/albedo/depth) into a G-buffer, then computes lighting per pixel from it → . The price of deferred is memory (the G-buffer is several full-screen images) and poor transparency. The modern compromise is Forward+ (tiled): the screen is split into 16×16 tiles and for each one you compute the list of lights that affect it — forward's transparency plus deferred's efficiency.
🕹 What to switch on — and what to notice
Render logistics are visible through engine debug views: overdraw, wireframe, the draw-call counter.
Almost every engine has an "overdraw" mode: the redder it is, the more times that pixel was repainted. Particles, smoke, foliage and UI layers glow red — that is where the fragment budget leaks away.
🎮 Do: in UE/Unity turn on the overdraw view and look at a scene with smoke or particles — you will see red zones of repeated repainting. Compare it with opaque geometry (almost no overdraw thanks to early-z). That is "don't paint a pixel twice" made visible.
In any open-world game, spin around 180° sharply and you will notice objects appearing as they enter the view frustum (frustum culling), and that what is behind a wall is not drawn (occlusion). LOD pop-in is the flip side of the same saving.
🎮 Do: in an open world, swing the camera quickly at the edge of the view distance — catch the moment things get drawn in. Step behind a large wall or building and ask yourself: is the engine drawing what is behind it? (It should not be — occlusion culling.) This is the work that does not get done.
Render statistics (Unity Stats, UE `stat rhi`, RenderDoc) show the number of draw calls and batches. A scene with thousands of unique objects is CPU-bound; the same scene with instancing is a few dozen calls.
🎮 Do: in the engine editor, look at the draw calls before and after enabling instancing/static batching on a field of grass or a crowd. Notice the call count dropping by orders of magnitude and the FPS rising — with the same image. The overhead was on the CPU, not the GPU.
Deep end · GPU architecture: warps, occupancy, divergence, coalescingskippable
Render optimization is working with the GPU architecture rather than against it.
Warps and divergence
Threads execute in groups of 32 (a warp, on NVIDIA) in lockstep — one instruction for all 32. If an if inside a warp splits threads across branches, the GPU executes both branches in sequence, masking the inactive threads — effective utilization halves (or worse with nested branches). Hence the rule: branch on data such that neighboring threads take the same branch (coherent branching).
Occupancy — how latency gets hidden
A global memory access takes hundreds of cycles. The GPU hides them not with a cache but by switching warps: while one waits on memory, another computes. The more active warps (occupancy), the more completely the latency is hidden. Too many registers or too much shared memory per thread lowers occupancy (fewer warps fit). Ray tracing tanks occupancy (long stalls traversing the BVH) — you need many threads in flight.
Memory coalescing
If thread 0 reads address X, thread 1 reads X+1, and so on, the GPU merges that into a single transaction (coalesced). Scattered access (strided/random) → dozens of transactions for the same volume. That is why a data-oriented layout (SoA, dense arrays) matters so much for GPU code: it makes access coherent. Shared memory (a per-block cache) is the manual way to reuse data without going out to global memory.
Deep end · engineering: explicit APIs, command buffers and tiled forward+skippable
Why explicit APIs
In OpenGL/D3D11 you set state and the driver guessed how to optimize it — and did so on one thread, becoming the bottleneck on multi-core CPUs. Vulkan/DX12/Metal/WebGPU made everything explicit: you assemble command buffers yourself and manage memory, descriptors and synchronization. The payoff is multi-threaded command recording (N threads → N buffers → submit at the end of the frame) and predictability; the price is more code and more ways to get it wrong. The pattern (command buffers / descriptor sets / barriers) carries between APIs — learn one and the rest are a different vocabulary.
Tiled Forward+
Deferred changes the complexity class of lighting from to , but pays in G-buffer memory and breaks transparency (only one depth layer). Forward+ splits the screen into tiles, uses a compute pass to build a list of affecting lights per tile (light culling), then forward-shades an object with only the relevant lights. You get both transparency (forward) and near-deferred efficiency. This is an example of changing the complexity class by restructuring rather than by micro-optimizing.
ML / AI (your domain): this is literally GPU utilization in training and inference. Warps/occupancy ⇄ tensor-core utilization and batch size (a small batch starves the GPU, like low occupancy); memory coalescing ⇄ contiguous/coalesced tensor access (and why layout decides things); divergence ⇄ why dense operations beat sparse ones on a GPU. "Don't send work that will be thrown away" = pruning / early exit / MoE routing / sparsity; "batch to amortize per-call overhead" = batching inference requests and amortizing kernel launches. Deferred (changing by restructuring) ⇄ an algorithmic change of complexity class (caching/memoization/linear attention). And explicit-API multi-threaded command submission ⇄ host→device as the bottleneck: often it is not the GPU that starves but the data loader / CPU-side feed (the same disease as the single-threaded driver).
Systems / data: predicate pushdown in a database is culling (don't read rows you will filter out); batching queries; staged pipelines with backpressure; latency hidden by parallelism (as with occupancy).
Performance engineering in general: the cheapest way to go faster is not doing the work (culling), then doing it in batches (batching), then changing the algorithm (complexity class), and only then micro-optimizing.
Principle: optimize in order of leverage — don't do unnecessary work, do it in batches, change the complexity, feed the parallel machine coherently. Micro-optimizing a shader or a kernel comes last, not first.
Why does culling matter more than optimizing the shader itself?
What is overdraw and why is it the "quiet" killer?
Why are draw calls a CPU cost rather than a GPU one?
Forward or deferred — when do you use which?
The GPU is "high-latency but high-throughput" — how is that both at once?
- "Real-Time Rendering" (Akenine-Möller et al.) — the chapters on the pipeline, culling, forward/deferred.
- "GPU Gems" / "GPU Pro" — practical techniques for culling, overdraw and optimization.
- RenderDoc plus its docs — taking a real frame apart (draw calls, overdraw, state).
- vkguide.dev / Vulkan Tutorial — how an explicit API builds command buffers.
- Module 8, "Rendering Pipeline" → "Forward vs Deferred", "GPU Architecture Basics", "Modern explicit graphics APIs" (
08-technical-deep-dives.md).