← Module 8/Audio and DSP
RU
Module 8 · Technical deep dive

Audio and DSP

Sound is the invisible half of game feel and at the same time hard real-time DSP: every few milliseconds the hardware asks for the next buffer of samples, and the audio thread has to fill it in time — otherwise you get a click. The deadline is tighter than the frame budget and far less forgiving.
~16 min🔬 DSP + 🎲 design
The gist in 30 seconds
Game audio is real-time DSP. Sound is a stream of samples (48 kHz = 48,000 amplitude values per second; Nyquist: the sampling rate must be ≥ 2× the highest audible frequency, ~20 kHz). The hardware consumes samples continuously and asks for the next buffer via a callback; the audio thread has to fill it before the deadline, otherwise you get a click/dropout (a 256-sample buffer @ 48 kHz = a 5.3 ms deadline). The engine builds a DSP graph: sources → effects (reverb/EQ/filters) → spatialization (HRTF for a 3D position in headphones) → the mix (sum without clipping) → output. The dynamics come from adaptive music (Hades: stem layers cross-faded by combat state) and procedural audio (synthesized footsteps → endless variation, little memory). All of it is packaged by middleware (FMOD/Wwise) as event-driven and parametric — the sound designer works without touching engine code. The shape — a streaming signal pipeline under a hard deadline — is the same as any real-time DSP or inference.

The mechanism: a signal under a hard deadline

Sound as samples

Audio is amplitude over time, sampled at 44.1/48 kHz: each sample is one amplitude value. A second of stereo @ 48 kHz is 96,000 numbers. The frequency ceiling is set by the Nyquist theorem:

fmax= R2=24 kHz (at R=48 kHz)

Humans do not hear above 20 kHz, so 44.1/48 kHz is enough; all the DSP works over this stream of numbers.

The callback and the buffer deadline

Audio hardware pulls samples continuously and periodically calls a callback: "give me the next N samples". The audio thread has to produce them before the buffer empties, otherwise the speaker plays garbage or silence → a click/dropout. This is hard real-time, and the timing is finer than the frame's:

latency= NR= 25648000≈5.3 ms

Buffer size is a trade-off between latency and safety: a small one (256 samples = 5.3 ms) gives low delay but a narrow deadline; a large one is safer but the sound "lags behind". And a miss is heard instantly: the eye barely notices a dropped frame, while a missed audio buffer is a sharp click. That is why audio lives on a separate high-priority thread with a fixed tiny deadline — it cannot share the main thread's floating timing.

The DSP graph

Sound is a signal graph processed buffer by buffer:

sources stream effects HRTF/3D mix out reverb/EQ/filter sum, no clipping every node processes buffer by buffer; the mix sums sources into buses and keeps the total from exceeding ±1.0 (clipping = distortion)

Sources (decoded samples/streams) → effect nodes (reverb, EQ, low-pass — each transforms the stream) → spatialization → the mixer (sums sources into buses with priorities and volumes, keeping the total from clipping past ±1.0) → output (stereo/5.1/7.1).

Spatialization: HRTF

How do you convey a source's 3D position? HRTF (Head-Related Transfer Function) is a set of measured filters describing how sound from a given direction is altered by the head and ears on its way to the eardrum (interaural differences in time and loudness plus frequency colouring). Convolving the stream with the HRTF response for the source's direction gives binaural stereo: in headphones the brain localizes "an explosion at (10,0,5)" around your head. On speakers you get surround panning instead.

Dynamics: adaptive music and procedural audio

All of it is packaged by middleware (FMOD/Wwise): the DSP graph, spatialization and adaptive music as event-driven, parametric authoring — the sound designer triggers events and turns parameters without engine code.

🕹 What to play — and what to notice

Audio has to be listened to in headphones — you cannot see spatialization and layers.

Hades adaptive music in layers

The benchmark for vertical layers: exploration → combat (drums/guitar kick in) → clear (they fall away). The transitions are seamless — it is a cross-fade of stems by combat state, not a track change.

🎮 Play: in Hades, listen to the music before a fight, during it and after — notice how layers are added and removed without a break (same tempo and key, only the density changes). That is adaptive music: not "we switched the song" but layers mixed in to match the gameplay state.

Hellblade: Senua's Sacrifice binaural HRTF

The famous binaural recording: the voices in Senua's head are spatialized around you. In headphones the brain localizes them behind and beside you — a direct demonstration of HRTF.

🎮 Play: Hellblade in headphones, mandatory — notice how the voices feel positional (behind you, to your left) rather than "in stereo". That is HRTF convolution turning direction into a binaural signal. Take the headphones off and the illusion collapses (HRTF assumes two independent ears).

A dropout/click a buffer underrun, live

In a loaded or badly optimized game the sound sometimes crackles or stutters — that is the audio thread failing to fill the buffer by the deadline (an underrun). You hear it instantly, unlike a drop in FPS.

🎮 Notice: catch the moment a game stutters and the sound crackles — that is a missed audio deadline. Compare it with visual lag: you can swallow a frame, you cannot swallow a click. That is why audio gets its own high-priority thread with a hard, tiny buffer deadline.

Deep end · DSP: sampling, the buffer deadline, mixing and convolutionskippable

Sampling and Nyquist

A continuous signal is sampled R times per second; by Nyquist, frequencies up to R/2 are exactly representable. Above that comes aliasing (a high frequency masquerading as a low one), so an anti-aliasing low-pass filter goes in front of the sampler. 48 kHz → a 24 kHz ceiling, with margin over hearing (~20 kHz). Bit depth (16/24) sets the dynamic range (the amplitude quantization step) — a direct analogue of fixed-point precision.

The real-time callback

The audio callback is hard real-time: inside it you cannot allocate memory, take locks or touch files — anything that might block for longer than the buffer deadline causes an underrun and a click. Hence the rules of audio programming: lock-free queues for talking to the game thread, pre-allocation, streaming from disk on a separate (non-audio) thread. The deadline is N/R — 5–10 ms — and a miss is audible rather than smoothed over.

Mixing, clipping and convolution

The mix is the sum of the source samples; if the sum goes past ±1.0 you get clipping (the peak is chopped off = harsh distortion), hence bus volumes, limiters and compressors. Reverb and HRTF are convolution of the stream with an impulse response (of a room or of an ear): output = input ∗ IR. Long IRs are computed by fast convolution via FFT (partitioned convolution) to make the deadline. The same math as in signal filtering generally.

Deep end · design: adaptive music, procedural audio, middlewareskippable

Horizontal vs vertical

Adaptive music comes in a vertical flavour (stem layers of one track fading in and out with intensity — Hades) and a horizontal one (reassembling sections: intro→loop→bridge→outro by state, joined on musical beats so the rhythm is not broken). The two are often combined. The key subtlety is transitions on the beat: a switch or cross-fade is timed to the bar, otherwise the seam is audible.

Procedural: the function instead of the output

A recorded sound is a fixed asset; a procedural one is a generator (synthesis from oscillators, noise, envelopes, physical models). Upsides: endless variation (no two footsteps or gunshots identical), tiny memory, parameterization by gameplay (speed, surface, material). The downside is realism, and it needs DSP synthesis expertise. This is exactly "store the function, not the materialized output" applied to sound.

Why middleware

FMOD/Wwise move the DSP graph, spatialization and adaptive logic into visual authoring for the sound designer: they assemble events, buses, parameters and transitions without code, while the engine merely triggers events and sends parameters. This is a division of labor (like a shader graph for an artist) — and the reason AAA almost never writes audio against bare APIs.

Analogy
An audio system is a cook at an endless conveyor. The speaker's belt carries away a tray of samples every few milliseconds no matter what; the audio thread has to have the next tray ready — otherwise an empty tray reaches your ears (a click). Unlike the "graphics cook" (who sometimes serves a dish late — a dropped frame, which the eye barely notices), the audio cook's lateness is heard instantly. So sound gets its own dedicated cook (thread) with a strict tiny deadline, and everything — effects, mixing, 3D — is done tray by tray on the belt, without stopping.
Why it matters
Audio is the invisible half of game feel (feedback, the weight of a hit, immersion), and technically it is hard real-time DSP with a deadline finer and less forgiving than the frame's: a click cuts the ear more than an FPS drop cuts the eye. Understanding it as a signal graph with a buffer callback, spatialization via HRTF convolution, and sound generated adaptively or procedurally is part of the engineering picture. And the shape (a streaming real-time signal pipeline under a hard deadline) generalizes to any signal-processing or streaming-inference system.
🔁 Beyond games — where this transfers
The lesson is streaming real-time signal processing under a hard deadline: the buffer callback, the DSP graph, convolution, sampling.

ML / AI (your domain): audio DSP is a form of streaming inference, especially audio ML. The buffer-callback deadline ⇄ the latency budget of streaming ASR/TTS: streaming TTS has to emit audio chunks before the buffer empties — the same underrun constraint (and directly relevant to TTS latency for LLM NPCs). Sampling/Nyquist ⇄ the basis of any signal ML, and "sound as discrete frames" ⇄ VQ-VAE audio tokens (an audio codebook — see discretization). HRTF convolution ⇄ convolutional primitives and spatial-audio ML. Adaptive music (state → layer choice) ⇄ controlled/conditional generation. Procedural audio (synthesis vs recording) ⇄ generative vs retrieval ("store the function, not the output"). And neural vocoders/RAVE/neural TTS are already a DSP graph with a learned node: the frontier where DSP meets ML. A dedicated audio thread with a tiny deadline ⇄ isolating the latency-critical inference path.

Signals / DSP: sampling, convolution/filters, FFT, real-time buffering — the shared basis of audio, radio, sensors and control.

Real-time systems: a hard-deadline callback with no blocking (lock-free, pre-allocation), streaming with backpressure, a separate priority thread for the latency-critical path.

Principle: in streaming processing under a deadline a miss is heard or seen at once — isolate the critical path onto its own thread, don't block in the callback, compute buffer by buffer.

🔧 Run it and poke at it — on your home machine
What to play is above (🕹). Here — getting hands on the signal.
🔧 Tinker (FMOD Studio / Web Audio) ~40 min
Download FMOD Studio (free) or work in the browser via the Web Audio API: build a simple DSP graph (source → reverb → output), change the buffer size and hear the latency↔dropouts trade-off. Make a mini adaptive track from two stems cross-faded by a parameter. Synthesize a procedural "footstep" (a sine impulse plus noise with an envelope) and tune it to a "surface".
🧪 Test (listening) ~15 min, headphones
Play Hellblade/Hades in headphones: localize the sources (HRTF) and hear the music layers (adaptive). Then in any game catch a buffer underrun (a click under load). State why the audio callback must not block (allocation/lock/file) and how that relates to TTS streaming.
Checklist: built a DSP graph and heard the buffer trade-off; made a cross-faded adaptive track and a procedural footstep; localized HRTF in headphones; connected the audio deadline to streaming inference.
Connections
foundation
Job systems — audio lives on its own high-priority real-time thread under a hard buffer deadline; communication with the game thread is lock-free.
foundation
Game feel — sound is the other half of "juice": a hit without audio feedback feels limp.
adjacent
LLM NPCs — streaming TTS runs into the same audio deadline (deliver a chunk before the buffer empties).
adjacent
Hardware constraints — procedural audio is "store the function, not the output"; bit depth is fixed-point precision.
Questions worth asking
Why does audio need its own thread, and why is the deadline tighter than the frame's?
Because a miss is heard instantly and sharply, while a visual one is barely noticed. A dropped frame is a micro-hitch the eye and brain smooth over; a dropped audio buffer means the speaker plays silence or garbage, which is a distinct click the ear catches at once. On top of that, audio is consumed continuously and evenly (the belt of samples never stops), whereas rendering can "catch up" after a dip. Audio cannot share the main thread's floating timing (where a frame is 12 ms one moment and 30 the next) — it needs an even stream of buffers hitting a hard deadline of N/R. Hence the dedicated high-priority thread and the ban on blocking inside the callback.
What exactly does HRTF do, and why does it only work in headphones?
HRTF encodes how sound from a specific direction is altered by your head and ears before it reaches the eardrum: it arrives at one ear slightly earlier and louder (interaural differences), and is colored differently in frequency by the shape of the outer ear. Convolving a mono source with the HRTF for its direction produces two different signals — one for the left ear, one for the right — and the brain reconstructs the position. It works in headphones because each ear gets exactly its own signal. On speakers the sound from the left also reaches the right ear (crosstalk), destroying the fine interaural differences — so there you settle for panning/surround, while true binaural requires isolated ears (headphones) or complex crosstalk cancellation.
Buffer size: why not just take a small one for low latency?
Because a small buffer means a narrow deadline: the audio thread has to fill it more often, and any spike in load (a complex mix, the OS scheduler not giving it time) causes an underrun → a click. A large buffer is safer (more slack) but adds delay: the sound trails the action (you press and the shot comes 20 ms later), which hurts game feel and especially rhythm games and VR. It is the classic trade-off: latency N/R against resilience to jitter. You pick the smallest buffer that reliably makes it on the target hardware (usually 128–512 samples) — like the smallest fixed timestep that does not fall into a spiral.
Procedural or recorded audio — which is better?
Different trade-offs, as always with "function vs output". Recorded gives maximum realism and control (the sound engineer hears the final result), but the assets are fixed: memory grows and repetition becomes noticeable ("the same footstep for the thousandth time"). Procedural gives endless variation (it never repeats), tiny memory and parameterization by gameplay (speed/surface/material change the synthesis on the fly), but realism is usually lower and it needs DSP synthesis skill. In practice it is a hybrid: the base is recorded, and variation and reactivity are added procedurally (randomizing pitch and sample start is the simplest "semi-procedural" trick against repetition). The choice depends on what matters more: the authenticity of a specific sound (recording) or variety/reactivity/memory (synthesis).
Why is adaptive music harder than just "switch the track in combat"?
Because an abrupt track change sounds like a seam and breaks immersion: it cuts off mid-phrase and jumps in key or tempo. Adaptive music makes the transition musically seamless: vertically, by mixing in layers of the same track (same key and tempo, only the density changes — Hades), or horizontally, by reassembling sections joined on a musical beat (waiting for the end of the bar). That requires the composer to write music in layers/sections for the system rather than linearly, and the engine to switch on the beat. The complexity pays for itself: the music breathes with the gameplay invisibly rather than clicking between modes.
Further reading