SNES Mode 7: pseudo-3D via an affine transform
1/z, which real 3D computes per pixel, is computed here once per line.(Sx,Sy) the hardware computes, by inverse mapping, which texel of the map to take the color from. On its own that is a flat rotate and zoom. The illusion of "a track receding to the horizon" comes from a per-scanline hack: HDMA changes the matrix on every scanline, and each line gets its own scale ∝ 1/p (p being the pixels below the horizon). The "3D objects" are ordinary sprites pasted on top. A cheap illusion cast in silicon.
The mechanism: the affine matrix and inverse mapping
An ordinary background is drawn tile by tile. Mode 7 instead takes a 1024×1024 square of map and stretches it onto the screen with an arbitrary affine transform — rotation, scale, translation, shear (everything except bending: straight lines stay straight). The hardware's key trick is that it goes not "where would each map texel land" (forward mapping would leave holes and overlaps) but the other way around: it walks the screen pixels left to right, top to bottom, and for each one asks "which texel am I from?". That is backward mapping — no holes, no overlaps.
For a screen pixel (Sx,Sy) the texel (Tx,Ty) is computed through the matrix [A B; C D] around the center point (Mx,My):
(plus the scroll M7HOFS/M7VOFS, which shifts the image after the transform). All coefficients are signed fixed-point 8.8: 8 bits of integer part, 8 fractional, divided by 256. The registers are named accordingly: M7A…M7D is the matrix, M7X/M7Y the center.
The geometric meaning of the matrix columns is direct: (A,C) is how many texels to step across the map when you move 1 pixel right on screen; (B,D) is the map step for 1 pixel down. Everything follows from that:
- Identity
(1,0)(0,1)→ 1 pixel = 1 texel → an ordinary background. - Scale is the length of the vectors:
(2,0)(0,2)→ 2 texels per pixel → the map is half the size (zoomed out). - Rotation by θ is the matrix
(cosθ, sinθ)(−sinθ, cosθ). - Shear is an uneven addition to one column.
Where the perspective comes from: per-scanline scaling
The matrix itself is flat — it rotates and stretches but creates no depth. Perspective "into the horizon" is a different matrix on every scanline, poured in via HDMA. Why per line? Because a horizontal line of a flat floor lies at a constant depth z. The perspective division 1/z doesn't change along the line → it is enough to compute the scale once per line.
With the camera at height h above the floor, the horizon projects to a fixed screen line. For a line p pixels below the horizon, the floor depth and the scale (texels per pixel) are:
(f is the focal length in pixels). At the horizon p→0 → the scale → ∞ (infinitely compressed, far away); at the bottom of the screen p is large → the scale is small (close, big). That 1/p dependence is what HDMA bakes in: each line gets its own M7A=M7D=scale(p).
A worked example. Let h = 256 texels. A line near the horizon at p=8 → a scale of 32 texels per pixel; p=64 → 4; the bottom line p=128 → 2. The same 256-texel landmark takes 256/32 = 8 px at the horizon and 256/2 = 128 px under your wheels. There is your depth gradient: the far end is compressed 16 times harder than the near one — because scale(8)/scale(128) = 128/8 = 16.
🕹 Games to play — and what to notice
One technique — "rotate and scale a single layer" — gives you everything from a flat spinning room to a race to the horizon. For each case: how it was done and what to play to see it with your hands.
An SNES launch title, with no coprocessor. One flat track layer + per-scanline HDMA scaling = a road to the horizon at 60 frames. The car is a sprite in the center that never rotates at all: the steering turns the whole world underneath it (a rotation matrix on the layer). Above the horizon is a separate raster sky, below it that one transformable layer. Single player: a full-screen per-line matrix eats almost the entire frame budget.
🎮 Play: in F-Zero, enter a long turn and watch the car — it stands still while the world rotates around it. Catch the horizon line with your eye: everything above it is static sky, everything below is the single affine layer.
Perspective isn't enough — you also have to compute the matrix coefficients fast every frame (projection, pitch). The 65816 is weak at that, so a math coprocessor, the DSP-1, went into the cartridge. Mario Kart (1992) is also split-screen: each half is its own independent Mode 7 layer, and the lost area is compensated with a map or a rear-view mirror. It grew out of a two-player F-Zero prototype.
🎮 Play: in two-player Super Mario Kart, notice that the top and bottom of the screen are two different Mode 7 planes. Drive off the track and you'll see the "ground" tilt and scale as a background over the sprites (that "background in front of the sprite" trick).
A Mode 7 floor for the overview map; the towns and the party are ordinary sprites composited on top. The airship's tilt and banking = rotating the plane under a motionless ship sprite.
🎮 Play: get into the airship in FF6 and turn — the world map rotates beneath you (Mode 7), but the towns and the ship itself stay flat sprites that merely scale. That gap between "a transformable floor" and "flat stickers" is what gives away 2.5D rather than 3D.
Rooms that rotate as a whole, and zooms — a demonstration of pure Mode 7 rotation of a single layer. There is no depth: you can't walk behind an object, everything lies in one plane that gets spun.
🎮 Play: in Castlevania IV, reach the rotating room — notice that you can't end up "behind" anything: it is one flat layer being spun about the center (M7X,M7Y).
Axelay's vertical stages imitate a "curved planet" with per-line raster scrolling (line scroll, the "rolling pin"), not Mode 7: the surface bends but doesn't rotate. The Sega Genesis had no hardware Mode 7 at all — the effect was faked in software, expensive for the CPU and worse-looking (Tales of Phantasia on the SNES; on the Genesis, isolated software tricks).
🎮 Play: compare Axelay's curved horizon (line scroll) with F-Zero's real Mode 7 floor. The test: line scroll bends lines but can't rotate the scene; an affine layer rotates, but its lines stay straight.
Deep end · theory: inverse mapping, fixed-point and why perspective is "free" per lineskippable
Why backward mapping specifically
Forward mapping (texel → screen pixel) leaves holes when magnifying (neighboring texels fly apart) and overlaps when minifying (several texels into one pixel, last writer wins, shimmering). Backward mapping (pixel → texel) guarantees exactly one read per output pixel: dense, unambiguous coverage. The price is that you have to be able to invert the transform; for an affine one that is just the inverse matrix, which is why the hardware computes (Tx,Ty) from (Sx,Sy) rather than the other way round.
Perspective as a degenerate case of 1/z
Perspective-correct texturing requires a division by depth for every screen pixel: u = (u/z)·z. In general 3D the depth z changes along a span → a division per pixel (expensive; Quake did it once every 16 pixels — see below). But a horizontal line of a flat floor under the camera lies at constant z → 1/z is constant along the line → one division per line. Mode 7 is honest perspective in the degenerate case of "the whole line at one depth". HDMA simply feeds the computed scale(p)=h/p line by line.
Fixed-point 8.8 and shimmering
The map step (A,C) is stored as 8.8 (a precision of 1/256 of a texel). Near the horizon the scale is enormous (tens of texels per pixel), so a tiny fractional difference in the step between frames jumps across many texels. With no mip levels and no filtering (Mode 7 has neither; sampling is nearest) this gives aliasing and "shimmer" in the distance. The same reason a distant texture without mipmaps sparkles on any hardware.
Deep end · engineering: HDMA, the frame budget and the coprocessorskippable
- HDMA = DMA by scanline. Before each scanline the DMA controller writes new
M7A…M7Dvalues from a table in RAM itself, without touching the CPU. That is how "a different matrix per line" is set without the processor's real-time involvement. - The CPU still computes the table. Only rasterization is "free". Somebody has to compute the
scale(p)coefficients for ~224 lines every frame — that is CPU work (or the DSP-1). Hence the1/ptable is often precomputed. - The DSP-1 is a math coprocessor on the cartridge (Pilotwings, Super Mario Kart): projection, matrix multiplication, trigonometry. A bare 65816 only manages the simple case of a single flat layer (F-Zero) — which is why F-Zero has no chip and Mario Kart does.
- One layer. Mode 7 is BG1 and only BG1; there are no parallax layers. The sky above the horizon is drawn separately (a raster split of color or a mid-frame settings change), and the "objects" are sprites on top. There is no real occlusion: you can "see through" a sprite, because there is no depth buffer.
- The Genesis: a faster CPU ≠ Mode 7. This is a fixed function in the SNES PPU, not a matter of megahertz. The Genesis VDP had no affine mode; software emulation ate the frame and looked worse.
Graphics / image processing: any texturing, image rotation or resize, image warping, lens remapping is a backward map plus interpolation (that is how holes are avoided). Quake's perspective-correct texturing is the general case of Mode 7's per-line division (the next lesson).
ML / AI: Spatial Transformer Networks and the grid_sample op are literally an inverse affine mapping of a feature map with bilinear sampling — Mode 7 in differentiable form; img2img warps and augmentations stand on the same thing. Deeper still is the principle of "hoist the invariant out of the inner loop": compute the expensive thing where it is constant (once per line) — that is loop-invariant code motion, KV-cache reuse (one prefix for many tokens), per-batch rather than per-example.
Performance / systems: a precomputed 1/p table instead of a per-pixel division is the classic "a table instead of arithmetic" (LUT). HDMA is a state change scheduled by scanline, the ancestor of the command buffer and per-draw-call uniform updates in modern rendering.
The principle: an approximation chosen to fit the structure of the problem (a plane → a line of constant depth) beats an honest general solution whenever that structure exists and you can see it.
M7A…M7D, the center and the scroll — you'll see rotation, scale and shear on a live texture. Then in bsnes-plus / Mesen-S, run F-Zero, open the layer viewer and trace HDMA: you'll see new scale values written into the matrix before every line. Turn BG1 off and the world vanishes, leaving the car sprite and the sky.If Mode 7 rotates one flat layer, why does F-Zero's car not turn while the world rotates around it?
Why is the perspective division once per line in Mode 7 but almost per pixel in Quake?
z → 1/z doesn't change along the line → one division per line. In real 3D a polygon span recedes, z changes along it → you need perspective correction per pixel. Quake amortized it by dividing once every 16 pixels and interpolating affinely in between. Mode 7 is the limiting case where "once every N pixels" became "once per whole line", because N = the full width.The Genesis had a higher CPU clock than the SNES — why does it have no Mode 7?
Why do Mario Kart and Pilotwings need the DSP-1 chip if Mode 7 is "free" in hardware?
Why does the texture near the horizon in Mode 7 shimmer and sparkle?
- SNESdev Wiki — "Mode 7 transform" and "Mode 7 perspective effects": the exact
M7A…M7D / M7X / M7Yregisters, the formula, HDMA. - Kulor, "Guide to Mode 7 Perspective Planes" (the NESDev forums) — how to build the
1/ptable. - Retro Game Mechanics Explained — "SNES Mode 7" (the mechanism explained in ~10 minutes).
- Module 2, the "SNES Graphics: Mode 7" section (
02-console-wars-1985-1993.md).