All terms

CRO Glossary

Simpson's Paradox

A statistical phenomenon where a trend appears in several groups of data but disappears or reverses when the groups are combined.

Simpson's paradox occurs when segment-level results point one direction, but the aggregated, overall result points the other way, because the mix of traffic across segments differs between variants. It's a classic trap in experimentation: a variant can look like a clear winner overall while actually being a loser (or a tie) within every single meaningful segment, once you slice by device, traffic source, or user type.

This usually happens because of unequal segment weighting between arms. For example, imagine Variant B performs slightly worse than Variant A on both mobile and desktop individually, but Variant B's traffic happened to include a higher proportion of desktop users (who convert better regardless of variant). The overall numbers can show Variant B "winning" purely due to this mix effect, not because the design change helped anyone.

Guarding against this means checking for sample ratio mismatch and comparing segment-level breakdowns, not just the topline number, before declaring a winner — especially in tests with multiple traffic sources or when running tests across different markets or devices simultaneously.

Related terms

See this in the wild

ABWatcher watches how top teams apply simpson's paradox.

Live A/B tests at 1,000+ high-converting brands, with plain-English hypothesis and takeaway.