Product Management Craft · 22 min · 150 XP
A/B tests you can trust
Run an A/B test you can trust: decide the sample size and the decision rule first, then don't peek, and check the split before believing the result.
The statistics of comparing groups (rates, not counts, and Simpson's paradox) are in Data Analytics Toolkit → Comparisons that hold up. This lesson is the product manager's side: setting a test up so its result means something, and the ways teams fool themselves when reading it.
An A/B test shows a random half of users the change and the other half the current version, then compares one metric. It's the most reliable way to know whether a change caused an improvement. It's also easy to run in a way that produces confident nonsense. Almost all of the mistakes happen before the test starts or while it's running, not in the maths at the end.
Write the plan before you start: the hypothesis, the primary metric and its guardrail, the smallest change worth detecting, and the decision you'll make for each result. The smallest worthwhile change decides the sample size: detecting a 1-point change in conversion takes far more users than detecting a 5-point one. Use a sample size calculator, and run for whole weeks, because Monday's users behave differently from Saturday's.
Don't peek and stop early. Results swing randomly at the start. If you check every day and stop the first time the result looks significant, you'll “find” winners that are really noise far more often than the 5% your significance level promised. Decide the end date, and read the result then. (Some testing tools use sequential methods designed for continuous monitoring; follow their rules, not your impatience.)
Planned split 50 / 50 Actual users A 51,840 B 48,160 (51.8% / 48.2%) With 100,000 users, a gap this big almost never happens by chance. Something is losing B users: a redirect, a crash, a bot filter. Don't read the result. Find the bug.
Two more checks. Sample ratio mismatch: if you planned a 50/50 split and got something noticeably different on a large sample, the assignment is broken and the result can't be trusted, whatever it says. And the novelty effect: people click on anything new for the first few days. A result driven entirely by week one often fades.
“Not significant” doesn't mean “no effect”. It means this test couldn't tell. If the test was too small to detect the change you care about, a flat result is inconclusive, not proof that the idea failed. That's why the sample size is decided from the smallest change worth detecting, up front.
Do it this week: for the next test your team runs, write the hypothesis, the metric, the guardrail, the smallest change worth detecting, the end date and the decision for each outcome on one page, before it starts.
Loading your workspace…