Analytics: telemetry, A/B and cohorts
The mechanism: from an event to a decision
Telemetry — the raw material
It all starts with event logging: session_start, level_complete(id, time), purchase(sku, $), churn. The stream (often millions of events/min) flows through a stream (Kafka-like) into storage and dashboards. This is the raw material: without instrumentation you are blind and decide by taste. The rule: log generously and early — you cannot analyze an event you never recorded.
Cohorts — why averages lie
Metrics are read by cohort — groups by install date (and UA source). "D7 for the March 1st cohort" is the truth; "average DAU" is a lie, because fresh installs mask the churn of older players: total DAU can grow while every cohort rots. Only a cohort shows the true shape of the retention curve and the true LTV. Splitting by source also catches "bad traffic" (cheap installs with zero retention).
Funnels — where it leaks
A funnel is the share passing each step of a sequence. An onboarding funnel, for example:
| Step | Remaining | Step conversion |
|---|---|---|
| Install | 1000 | — |
| Finished the tutorial | 600 | 60% |
| Reached level 5 | 300 | 50% |
| First purchase | 30 | 10% |
You fix the worst step (the biggest drop relative to expectation) — here the "install→tutorial" cliff (−40%) is often the most expensive, because it hits D1. A funnel turns "retention is bad" into "here is exactly where we lose them".
A/B — testing causality
A hypothesis is not "discussed" but tested: players are split randomly into control and variant, one thing changes (a button color, offer timing, the difficulty curve), a metric is measured and the winner is shipped — but only at statistical significance, otherwise you ship noise. The key math is that the required sample size grows as the inverse square of the effect:
where is the detectable effect (MDE) and is the spread. Want to catch a lift half the size — you need four times the users. The main traps: peeking (checking a running test and stopping it as soon as it looks "significant" — this inflates false positives from 5% to ~15–30%), multiple comparisons (test 20 metrics and one is "significant" by chance), the novelty effect (anything new gets a temporary boost) and the need for guardrail metrics (do not win on retention by collapsing revenue).
The cycle and the kill-fast culture
All of it runs as a loop: instrument → observe (cohorts/funnels) → hypothesize → A/B → ship or kill → repeat. Supercell turned this into a culture: small autonomous "cell" teams, games soft-launched in a couple of countries, and if the metrics (D1/D7, early LTV) miss the thresholds the game is killed fast, without sentiment. Most of their games die in soft launch and never ship globally; a few survive (Clash Royale, Brawl Stars). Data here is not reporting but a selection mechanism.
🕹 What to open — and what to notice
What you "play" here is the dashboard: the data is visible both in public reports and in the way a game instruments you.
From your first launch you are in a funnel and, most likely, in an A/B bucket: the timing of the first offer, the length of the tutorial, the first reward — all of it is instrumented and possibly being tested on you right now.
🎮 Notice: install a new F2P game and track your own onboarding funnel — where you are gently nudged (a first "win" in the first seconds, for D1), when the first offer appears. Work out what you would A/B test if you were the producer.
Open reports show real retention curves and funnels by genre — the numbers teams line their own cohorts up against (median D1 ~23%, D7 ~4%).
🎮 Watch: open a recent mobile benchmark report and read the cohort curve and the distribution across top / median / bottom 25%. Notice how different "the average" is from "by cohort" — and why teams look at the second one.
Their games ship first in selected countries, and the data decides their fate. Dozens of projects were shut down in soft launch (Smash Land, Spooky Pop and others) and only a few survived — data as a filter, not as a report.
🎮 Watch: read through the list of killed Supercell games and the thresholds they were closed on. Notice the principle: better to kill fast on early cohorts than to keep going on taste. It is the same discipline as the vertical slice in indie production.
Deep end · statistics: sample size, peeking and multiple comparisonsskippable
An A/B test is hypothesis testing, and every classic statistical trap is real here and expensive.
Sample size and power
To reliably detect an effect at variance you need per arm (with a constant set by your chosen α and power). Small effects (and in a mature game the lifts are fractions of a percent) demand enormous samples: half the effect → four times the users and time. Which is why small games cannot A/B test everything — there is not enough traffic to reach significance.
Peeking — the most common mistake
Checking a running test and stopping the moment you see "p<0.05" inflates false positives from 5% to ~15–30%: under repeated checks almost any noise will cross the threshold eventually. The cure is either a sample size fixed in advance or sequential tests (alpha spending, always-valid confidence intervals).
Multiple comparisons and guardrails
Test 20 metrics and on average one is "significant" by chance (you need a Bonferroni / FDR correction). And guardrail metrics are mandatory: a variant can raise the target (offer conversion, say) while collapsing a side metric (retention, refunds) — without protective metrics you "win" into the red. And locality: A/B finds incremental improvements around the current design, not qualitative jumps — for those you need a new hypothesis, not a test.
Deep end · infra: the telemetry pipeline and real time versus batchskippable
Underneath analytics is a data pipeline: an event on the client/server → a stream (Kafka/Kinesis) → storage (a warehouse/lake) → dashboards/ML. Two modes:
- Real time (stream processing): live dashboards, "revenue dropped" alerts, anti-cheat, dynamic offers. Expensive and complex, and needed for operational decisions.
- Batch (nightly jobs): cohort reports, LTV models, heavy analytics. Cheaper, and no instant response required.
The key engineering pains: the event schema (change the format and you break historical reports; you need versioning), idempotency/dedup (events arrive twice), and cheap writes versus completeness (logging everything is expensive in bandwidth and storage — it is a balance). Authority sits on the server: critical events (purchases) are computed server-side and client numbers are not trusted (anti-cheat, see the authoritative server).
ML / AI (your domain): this is literally online experimentation and model evaluation. A/B = shipping a model behind a flag and measuring the lift; peeking ⇄ overfitting to the validation set (peek at the test set often enough and you select noise); multiple comparisons ⇄ model selection across many metrics (you need a correction or a holdout); Goodhart ⇄ reward hacking (optimize a proxy metric and the model breaks the spirit of the task); guardrail metrics ⇄ secondary constraints you are not allowed to collapse. Cohorts = the right way to evaluate any rollout; kill-fast soft launch = managing a portfolio of research bets on early signals. And the thread running through this course: a metric is a proxy for the goal, not the goal; knowing where a number misleads (data-driven versus judgment) is part of professionalism.
Product / growth: the entire growth stack — funnels, cohort retention, A/B platforms (Optimizely-like), a north-star metric; the same discipline of significance and guardrails.
Science / any experiment: power, sample size , p-hacking, pre-registration versus fitting after the fact — the same rules as in a live game's A/B.
The principle: measure causally (randomization + cohorts), respect the statistics (sample size, no peeking, protect your side metrics) — and remember you are optimizing a proxy, not the goal itself.
Why cohorts rather than "average DAU" — the average is simpler?
What is wrong with stopping an A/B as soon as you see significance?
Can A/B testing produce a qualitative leap in a game?
Goodhart: how does a "good metric" lead to harm?
Why do most Supercell games die in soft launch — is that a failure?
- Kohavi, Tang, Xu, "Trustworthy Online Controlled Experiments" — the A/B bible (peeking, guardrails, traps).
- GameAnalytics / deltaDNA — the practice of game telemetry, cohorts and funnels.
- Breakdowns of Supercell's cell structure and kill-fast (GDC talks, Ilkka Paananen interviews).
- "A/B Testing Infrastructure at Scale" — module 6 plus GDC talks on live analytics.
- Module 6, "A/B Testing Infrastructure" + "Real-Time Analytics Pipelines" + "Supercell's Cell Structure" (
06-mobile-f2p-live-service-2012-2018.md).