Test Enough and Something Wins

You tested twenty things and one worked. Did it?

The idea

When many comparisons are made, some will look significant by chance alone.

In the real world

Testing twenty variants at the usual threshold produces about one false winner.

Going deeper

At a 5% threshold, roughly one in twenty comparisons will look significant with no real effect. Twenty variants producing one winner is exactly what pure chance delivers.

This is a large part of why wins from experimentation fail to reproduce. The fix is either a stricter threshold adjusted for the number of tests, or replication on fresh data before acting. The reporting habit that prevents most of the damage is simply stating how many things were tested alongside the one that worked, since that number is what determines how surprised to be.

Where it stops applying

Corrections for multiple comparisons reduce power and can hide real effects, which has its own cost. The aim is honest accounting of how many chances were taken.

Why it matters

It explains why so many wins from experimentation fail to reproduce.

Try this today

Write down how many things you tested before reporting the one that worked.

Test yourself

A team tests twenty button variants and one is significant at p<0.05. Why is that not yet a finding?

Show the answer

At that threshold roughly one in twenty will look significant with no real effect, so a single winner from twenty tests is exactly what pure chance produces. The result needs either a stricter threshold for the number of tests, or replication on fresh data.

Learn this in the feed Answering from memory, then again days later, is what makes it stick.

More in Data & Statistics

The Sample Decides the Answer Variance Matters as Much Relative Without Absolute Proxies Drift Axes Can Lie Regression to the Mean

All Data & Statistics lessons