A/B Testing
A/B testing shows two versions of something to different users at random, then compares a chosen metric to see which performed better.
Because the split is random, a difference big enough to be unlikely by chance can be credited to the change itself. That causal claim is the whole value. Analytics can show two things moved together, but only an experiment supports saying one caused the other.
How it works
Users are randomly assigned to version A or version B, and the assignment sticks so nobody sees both. One metric is chosen in advance. At the end, the metric is compared between groups.
The conventional bar is a p-value under 0.05, meaning a result this extreme would happen by chance less than one time in twenty if there were no real difference.
Sample size is decided before starting. This matters more than it sounds, and the reason is in the trade-offs below.
Trade-offs
- Significance is not size. With enough traffic, a difference too small to care about becomes statistically significant. Whether it is worth shipping is a separate judgement.
- It tells you what, never why. A variant can lift signups while quietly damaging trust, and the test will report a win because trust was not the metric.
- It needs traffic. Detecting a small improvement takes far more visitors than most products have. Low volume is exactly where convincing false positives come from.
Common mistakes
- Peeking. Checking repeatedly and stopping the moment the line crosses the threshold will manufacture significance out of noise. The maths assumes one check at a predetermined sample size.
- Reading a null result as proof. No significant difference may simply mean the test was too small to detect one.
- Testing without a hypothesis. Running variants to see what sticks produces winners that do not repeat.
Key takeaways
- Fix the sample size before starting, then run to it.
- Significant and meaningful are different questions.
- A win on your chosen metric may be a loss on one you did not measure.
- Low traffic makes false positives more likely, not less.
Learn this
Lessons and exercises mapped to this concept.
Common questions
- How is A/B testing different from usability testing?
- A/B testing measures what happens at scale with statistical confidence and no explanation. Usability testing observes why it happens, in depth, with a handful of people and no statistical claim. A/B testing tells you version B converts better. Only usability testing tells you version A confused people about the price.
- Why can I not stop the test as soon as it looks significant?
- Because the numbers move around throughout the test, and if you keep checking you will eventually catch a random swing crossing the line. Stopping there records that noise as a finding. The significance calculation assumes a single check at a sample size you committed to in advance.
- Can a small product run useful A/B tests?
- Often not, and it is worth admitting that. Detecting a modest improvement needs far more traffic than most small products get. With limited volume you usually learn more from talking to users, and from only testing changes large enough to produce an obvious effect.