Audio and DSP
The mechanism: a signal under a hard deadline
Sound as samples
Audio is amplitude over time, sampled at 44.1/48 kHz: each sample is one amplitude value. A second of stereo @ 48 kHz is 96,000 numbers. The frequency ceiling is set by the Nyquist theorem:
Humans do not hear above 20 kHz, so 44.1/48 kHz is enough; all the DSP works over this stream of numbers.
The callback and the buffer deadline
Audio hardware pulls samples continuously and periodically calls a callback: "give me the next N samples". The audio thread has to produce them before the buffer empties, otherwise the speaker plays garbage or silence → a click/dropout. This is hard real-time, and the timing is finer than the frame's:
Buffer size is a trade-off between latency and safety: a small one (256 samples = 5.3 ms) gives low delay but a narrow deadline; a large one is safer but the sound "lags behind". And a miss is heard instantly: the eye barely notices a dropped frame, while a missed audio buffer is a sharp click. That is why audio lives on a separate high-priority thread with a fixed tiny deadline — it cannot share the main thread's floating timing.
The DSP graph
Sound is a signal graph processed buffer by buffer:
Sources (decoded samples/streams) → effect nodes (reverb, EQ, low-pass — each transforms the stream) → spatialization → the mixer (sums sources into buses with priorities and volumes, keeping the total from clipping past ±1.0) → output (stereo/5.1/7.1).
Spatialization: HRTF
How do you convey a source's 3D position? HRTF (Head-Related Transfer Function) is a set of measured filters describing how sound from a given direction is altered by the head and ears on its way to the eardrum (interaural differences in time and loudness plus frequency colouring). Convolving the stream with the HRTF response for the source's direction gives binaural stereo: in headphones the brain localizes "an explosion at (10,0,5)" around your head. On speakers you get surround panning instead.
Dynamics: adaptive music and procedural audio
- Adaptive music — it reacts to gameplay. Hades: a base layer (exploration) plus a combat layer (drums kick in during a fight) plus a boss layer (guitar). Stems cross-fade by state (vertical layers) or sections get reassembled (horizontally). Parameters (intensity, tension) drive the mix.
- Procedural audio — synthesis instead of recording: a footstep is an impulse sine plus noise with an envelope, tuned to the surface (dirt/metal/stone). The upside is endless variation and little memory; the downside is less realism and the need for synthesis skill ("store the function, not the output").
All of it is packaged by middleware (FMOD/Wwise): the DSP graph, spatialization and adaptive music as event-driven, parametric authoring — the sound designer triggers events and turns parameters without engine code.
🕹 What to play — and what to notice
Audio has to be listened to in headphones — you cannot see spatialization and layers.
The benchmark for vertical layers: exploration → combat (drums/guitar kick in) → clear (they fall away). The transitions are seamless — it is a cross-fade of stems by combat state, not a track change.
🎮 Play: in Hades, listen to the music before a fight, during it and after — notice how layers are added and removed without a break (same tempo and key, only the density changes). That is adaptive music: not "we switched the song" but layers mixed in to match the gameplay state.
The famous binaural recording: the voices in Senua's head are spatialized around you. In headphones the brain localizes them behind and beside you — a direct demonstration of HRTF.
🎮 Play: Hellblade in headphones, mandatory — notice how the voices feel positional (behind you, to your left) rather than "in stereo". That is HRTF convolution turning direction into a binaural signal. Take the headphones off and the illusion collapses (HRTF assumes two independent ears).
In a loaded or badly optimized game the sound sometimes crackles or stutters — that is the audio thread failing to fill the buffer by the deadline (an underrun). You hear it instantly, unlike a drop in FPS.
🎮 Notice: catch the moment a game stutters and the sound crackles — that is a missed audio deadline. Compare it with visual lag: you can swallow a frame, you cannot swallow a click. That is why audio gets its own high-priority thread with a hard, tiny buffer deadline.
Deep end · DSP: sampling, the buffer deadline, mixing and convolutionskippable
Sampling and Nyquist
A continuous signal is sampled times per second; by Nyquist, frequencies up to are exactly representable. Above that comes aliasing (a high frequency masquerading as a low one), so an anti-aliasing low-pass filter goes in front of the sampler. 48 kHz → a 24 kHz ceiling, with margin over hearing (~20 kHz). Bit depth (16/24) sets the dynamic range (the amplitude quantization step) — a direct analogue of fixed-point precision.
The real-time callback
The audio callback is hard real-time: inside it you cannot allocate memory, take locks or touch files — anything that might block for longer than the buffer deadline causes an underrun and a click. Hence the rules of audio programming: lock-free queues for talking to the game thread, pre-allocation, streaming from disk on a separate (non-audio) thread. The deadline is — 5–10 ms — and a miss is audible rather than smoothed over.
Mixing, clipping and convolution
The mix is the sum of the source samples; if the sum goes past ±1.0 you get clipping (the peak is chopped off = harsh distortion), hence bus volumes, limiters and compressors. Reverb and HRTF are convolution of the stream with an impulse response (of a room or of an ear): output = input ∗ IR. Long IRs are computed by fast convolution via FFT (partitioned convolution) to make the deadline. The same math as in signal filtering generally.
Deep end · design: adaptive music, procedural audio, middlewareskippable
Horizontal vs vertical
Adaptive music comes in a vertical flavour (stem layers of one track fading in and out with intensity — Hades) and a horizontal one (reassembling sections: intro→loop→bridge→outro by state, joined on musical beats so the rhythm is not broken). The two are often combined. The key subtlety is transitions on the beat: a switch or cross-fade is timed to the bar, otherwise the seam is audible.
Procedural: the function instead of the output
A recorded sound is a fixed asset; a procedural one is a generator (synthesis from oscillators, noise, envelopes, physical models). Upsides: endless variation (no two footsteps or gunshots identical), tiny memory, parameterization by gameplay (speed, surface, material). The downside is realism, and it needs DSP synthesis expertise. This is exactly "store the function, not the materialized output" applied to sound.
Why middleware
FMOD/Wwise move the DSP graph, spatialization and adaptive logic into visual authoring for the sound designer: they assemble events, buses, parameters and transitions without code, while the engine merely triggers events and sends parameters. This is a division of labor (like a shader graph for an artist) — and the reason AAA almost never writes audio against bare APIs.
ML / AI (your domain): audio DSP is a form of streaming inference, especially audio ML. The buffer-callback deadline ⇄ the latency budget of streaming ASR/TTS: streaming TTS has to emit audio chunks before the buffer empties — the same underrun constraint (and directly relevant to TTS latency for LLM NPCs). Sampling/Nyquist ⇄ the basis of any signal ML, and "sound as discrete frames" ⇄ VQ-VAE audio tokens (an audio codebook — see discretization). HRTF convolution ⇄ convolutional primitives and spatial-audio ML. Adaptive music (state → layer choice) ⇄ controlled/conditional generation. Procedural audio (synthesis vs recording) ⇄ generative vs retrieval ("store the function, not the output"). And neural vocoders/RAVE/neural TTS are already a DSP graph with a learned node: the frontier where DSP meets ML. A dedicated audio thread with a tiny deadline ⇄ isolating the latency-critical inference path.
Signals / DSP: sampling, convolution/filters, FFT, real-time buffering — the shared basis of audio, radio, sensors and control.
Real-time systems: a hard-deadline callback with no blocking (lock-free, pre-allocation), streaming with backpressure, a separate priority thread for the latency-critical path.
Principle: in streaming processing under a deadline a miss is heard or seen at once — isolate the critical path onto its own thread, don't block in the callback, compute buffer by buffer.
Why does audio need its own thread, and why is the deadline tighter than the frame's?
What exactly does HRTF do, and why does it only work in headphones?
Buffer size: why not just take a small one for low latency?
Procedural or recorded audio — which is better?
Why is adaptive music harder than just "switch the track in combat"?
- "Game Audio Implementation" (Richard Stevens, Dave Raybould) — FMOD/Wwise and adaptive music in practice.
- FMOD Studio / Wwise docs plus the free versions — the DSP graph, events and parameters by hand.
- "The Audio Programming Book" / Will Pirkle, "Designing Audio Effect Plugins" — DSP, convolution, the real-time callback.
- GDC audio talks — Hellblade (binaural), Hades (adaptive), procedural audio (No Man's Sky).
- Module 8, "FMOD/Wwise Architecture", "Spatial Audio & HRTF", "Dynamic Music", "Procedural Audio" (
08-technical-deep-dives.md).