All terms

CRO Glossary

Confidence Washing

Presenting an underpowered or flawed test result with false statistical authority to justify a decision already made.

Confidence washing happens when a team wraps a weak or inconclusive experiment in the language of rigor — p-values, confidence intervals, dashboards — to make a predetermined decision look data-driven. The numbers are real, but the process behind them was bent to fit a conclusion: the test was stopped early because it looked good, a segment was cherry-picked because it hit significance, or a low-powered test was reported as a clear win because stakeholders wanted to ship.

This matters because it erodes trust in the whole experimentation program. Once a team discovers that a 'statistically significant' result was actually the product of peeking, selective segment reporting, or a suspiciously convenient stopping rule, they start questioning every past result, not just the flawed one. It's a cultural problem as much as a statistical one — usually driven by pressure to ship, not by dishonesty.

A typical example: a redesign test runs for three days, shows a positive trend in one browser segment, and gets reported to leadership as 'the new checkout increases conversion by 12%, with statistical significance,' omitting that the overall result was flat and the sample size was a fraction of what the pre-registered plan called for.

The fix is procedural: pre-register hypotheses and stopping rules, report overall results alongside any segment cuts, and require analysts to flag when a result technically clears a significance threshold but fails other quality checks (power, duration, SRM).

Related terms

See this in the wild

ABWatcher watches how top teams apply confidence washing.

Live A/B tests at 1,000+ high-converting brands, with plain-English hypothesis and takeaway.