All terms

CRO Glossary

Statistical Power

The probability that a test correctly detects a true effect when one actually exists.

Statistical power is typically expressed as a percentage, commonly 80% or 90%, and represents how likely your experiment is to detect a real difference between variants if that difference truly exists. Low power means a test might run to completion and show no significant result even though the treatment genuinely works — a false negative, not proof of no effect.

Power depends on three things: sample size, the size of the effect you're trying to detect (MDE), and the variability of the metric being measured. Increasing sample size or accepting a larger MDE both raise power; noisy metrics with high variance require more traffic to reach the same power level.

A practical example: if a team runs a test with too few visitors and finds no significant lift, it's tempting to conclude the change 'doesn't work.' But if the test was underpowered — say, only 40% power — there was a good chance of missing a real effect entirely. Calculating required sample size for adequate power before launching a test is one of the most commonly skipped, and most costly to skip, steps in experimentation programs.

Related terms

See this in the wild

ABWatcher watches how top teams apply statistical power.

Live A/B tests at 1,000+ high-converting brands, with plain-English hypothesis and takeaway.