Skip to content
← Blog

Why Most A/B Tests Are Lying to You

Most A/B tests never reach statistical validity before someone calls a winner. Here's how to run tests that actually mean something.

Someone on the team runs an A/B test on the signup button. Three days in, variant B is up 12%. Excited, they call it, ship the winner, and add a line to the next all-hands about the lift. Two months later, signups haven’t moved. Nobody goes back to check why, because by then there are five other tests to report on.

This happens constantly, and it isn’t bad luck. It’s what happens when a statistical method gets used as a decision-making ritual instead of a statistical method.

The stopping problem

The single biggest source of false positives in A/B testing is peeking. A test is designed to reach a conclusion at a predetermined sample size, with a predetermined significance threshold. Checking the dashboard every day and stopping the moment the p-value crosses 0.05 is not the same test. It’s a different procedure entirely, and it inflates your false positive rate dramatically — often to 20-30% instead of the 5% everyone assumes they’re getting.

The math is straightforward: with continuous monitoring, you get many chances for random noise to cross the significance line at some point during the run, even when there is no real effect. Stop the first time it crosses, and you’ve essentially guaranteed you’ll “find” winners that aren’t real a meaningful fraction of the time.

The fix isn’t complicated, just unpopular: decide your sample size and test duration before you start, based on a power calculation, and don’t look at results as a stopping signal until you get there. If your organization can’t tolerate waiting three weeks for a real answer, that’s useful information — it means A/B testing isn’t the right tool for the decision you’re trying to make.

Underpowered tests produce noise, not signal

Most SMB and mid-market products don’t have the traffic to detect small effects quickly. If your signup page gets 2,000 visitors a month and you’re trying to detect a 5% lift in a 3% baseline conversion rate, you need a sample size that would take the better part of a year to accumulate at standard significance and power thresholds. Nobody runs that test for a year. They run it for two weeks and call whatever comes out significant.

The result is a stream of “wins” that are really just sampling variance dressed up as insight. Before you set up a test, run the power calculation. If the required sample size is unreachable in a reasonable window, don’t run a formal test — either batch several changes into a larger redesign you evaluate with before/after data, or accept that you’re making a qualitative call and stop pretending a p-value backs it up.

Multiple comparisons quietly break everything

Testing five variants against a control, or slicing results by device, browser, traffic source, and new-vs-returning, multiplies your chances of finding something “significant” by chance. Run 20 comparisons at a 5% significance threshold and you should expect roughly one false positive even if nothing you tested does anything at all. Teams that segment results after the fact — “it didn’t win overall, but look at mobile Safari users” — are almost always looking at noise, not a real subgroup effect.

If you want to test subgroups, decide which ones matter before you look at the data, and adjust your significance threshold accordingly. Post-hoc slicing until something turns green is not analysis. It’s a way of guaranteeing you’ll find whatever you’re looking for.

What actually correlates with a test being real

A few practical checks are more useful than the p-value alone:

The bigger point

A/B testing is a genuinely powerful tool when it’s used the way it was designed to be used: pre-registered hypothesis, calculated sample size, fixed stopping rule, one primary metric. It becomes theater when it’s used as a way to add statistical legitimacy to a decision the team already wanted to make. The tell is usually speed — real tests take longer than anyone wants them to, because reality doesn’t move as fast as a dashboard refresh.

If your testing program produces a steady stream of small wins that never seem to add up to the growth they implied, the problem probably isn’t your ideas. It’s the test.

PNK WORKS builds the analytics and experimentation infrastructure to run tests that hold up, not just ones that look good in a slide. Talk to us.

Ready to work together?

Start a Project →