A/B Test Sample Size Calculator: How Many Users You Actually Need
A step-by-step guide to calculating sample size, power, and significance for A/B tests — so your first experiment produces a real answer, not noise.
Maya Patel
Senior CRO Strategist · Jul 6, 2026
Most failed A/B tests don't fail because the idea was bad. They fail because someone shipped a test, waited two weeks, saw a "winner," and called it before the math caught up. In our monitoring of live experiments across 1,000+ high-traffic sites, the single most common root cause of a retracted or reversed result isn't a broken hypothesis — it's an underpowered sample that got called early.
If you're a growth PM running your first test, the sample size question isn't a formality to check off before launch. It's the difference between a decision backed by evidence and a coin flip dressed up in a confidence interval. Here's how to size it correctly, with the actual mechanics — not just "use a calculator."
Why Sample Size Is the Test, Not a Detail Around It
Every A/B test is a tradeoff between four inputs, and you only control three of them:
- Baseline conversion rate — what your control currently does (you measure this, you don't choose it)
- Minimum detectable effect (MDE) — the smallest lift you actually care about catching
- Statistical significance — your tolerance for false positives (industry standard: 95%, or p < 0.05)
- Statistical power — your tolerance for false negatives (industry standard: 80%)
Sample size is what falls out the other end once you fix the last three. Teams that skip this step and just "run it until it feels done" are implicitly accepting an unknown, usually terrible, power level — often below 50%. That means a coin-flip's chance of missing a real effect even when one exists.
Nielsen Norman Group's walkthrough frames this well with a concrete example: a 3% baseline CTA click rate, a 20% relative MDE (meaning you want to detect a shift to roughly 3.6%), at 95% significance. Plug those into a sample size formula and you get a number — not a vibe.
The Four Inputs, and How to Set Each One Realistically
Baseline rate: Pull this from your last 30-90 days of analytics on the exact metric and exact page/flow you're testing. Don't use an annual average if the metric is seasonal — a Q4 ecommerce checkout rate is not your July baseline.
MDE: This is where most first-time testers get overambitious. It's tempting to size for a 5% lift because that's what the roadmap deck says you need. But smaller MDEs require exponentially larger samples. The relationship isn't linear — halving your MDE roughly quadruples your required sample size. Per Analytics Toolkit's research, detecting a 5-10% relative effect at 95%/80% typically requires between 50,000 and 500,000 users, depending on your baseline conversion rate. If your traffic can't support that in a reasonable window, you either need a bigger MDE (only look for bigger lifts) or a different metric closer to the top of the funnel with more volume.
Significance (α): 95% is the default because it caps your false-positive rate at 5%. Some fintech and healthcare teams tighten this to 99% given the cost of a bad ship decision — that's a legitimate tradeoff, but it further inflates sample size, so know what you're paying for.
Power (1-β): 80% is standard, meaning you accept a 20% chance of missing a real effect that exists. Teams running high-stakes redesigns sometimes push to 90%, which again raises the sample requirement — roughly 30% more traffic needed to go from 80% to 90% power at the same MDE.
Worked Example: From Inputs to a Number
Say you're testing a new checkout CTA. Baseline conversion is 4%. You care about detecting a 15% relative lift (4% → 4.6%). You want standard 95% significance and 80% power.
Using a two-proportion z-test sample size calculator (any standard A/B tool — Evan Miller's, AB Tasty's, or the one referenced in this walkthrough), that combination lands you around 30,000+ visitors per variation. That's not a typo, and it's consistent with the benchmark Cro Metrics cites — roughly 30,244 visitors per variation to reliably detect a 10%+ effect on a similarly modest baseline.
Compare that to a simpler example from an email subject-line test with a much higher baseline open rate and a larger MDE: the required sample dropped to just 600 users per variation. The lesson isn't "sample size is always huge" — it's that sample size is entirely dictated by your specific baseline and MDE. There's no universal number, which is exactly why generic advice like "run it for two weeks" is meaningless without doing the calculation first.
Translating Sample Size Into Test Duration
Once you have your required sample size per variation, duration is simple division:
Duration (days) = (Sample size per variation × number of variations) / average daily eligible traffic
Two things trip teams up here:
Weekly seasonality. If your traffic mix shifts on weekends (common in B2C ecommerce) or Mondays (common in B2B SaaS), you need to run in full-week increments — minimum 7, ideally 14 days — even if you technically hit your sample number mid-week. Cutting a test on day 4 because you cleared the number introduces a day-of-week confound that will bias your read.
Traffic that doesn't scale to your number in a reasonable window. If the math says you need 45 days and your roadmap has a 2-week cycle, that's a signal — not a reason to under-power the test anyway. Options: broaden the audience, move the test higher in the funnel where volume is larger, or accept a larger MDE and only look for bigger wins.
Common Mistakes That Quietly Invalidate Results
Peeking and stopping early. Checking your dashboard daily and calling the test the moment it crosses significance inflates your false-positive rate dramatically — repeated peeking without a sequential testing correction can push your real error rate well above the 5% you think you're protecting. If you need to peek, use a testing platform with proper sequential/Bayesian correction built in, not a fixed-horizon calculator you're checking early.
Ignoring the "3,000 conversions" heuristic. Beyond raw visitor counts, GuessTheTest's guidance suggests aiming for at least 3,000 conversions per variation as a practical floor for result stability, on top of the 30,000-visitor benchmark. Low-conversion pages can hit visitor targets while still being conversion-starved and noisy.
Testing too many variations at once. Each additional variation splits your traffic and multiplies your comparisons, which either extends duration or forces you to accept lower power. A 4-way test needs meaningfully more total traffic than an A/B — budget for it in the sizing math, not after launch when velocity slows.
Confusing statistical significance with practical significance. A test can hit p < 0.05 on a 0.3 percentage-point lift that isn't worth the engineering cost to ship. Size your MDE around a lift that actually matters to the business, not the smallest number your calculator will accept.
Applying an ICE Lens Before You Size Anything
Before running the sample size math at all, run the hypothesis through Impact, Confidence, and Ease scoring. A high-Impact idea with weak Confidence (you're not sure the effect is even real) is exactly the case where proper sizing matters most — you don't want to under-power a speculative bet and get a false "no effect" read that kills a good idea prematurely. Conversely, a low-Ease test (heavy engineering lift) sized for a huge sample is a good candidate to de-scope the MDE or find a higher-traffic proxy metric before committing the sprint.
Your Next Sprint
Before you write a single line of test copy, do the sizing math first: pull your real baseline, decide the smallest lift worth caring about, pick your significance and power (95%/80% unless you have a specific reason to move), and calculate sample size per variation. Then convert that to a duration in full weeks against your actual daily traffic. If the number doesn't fit your roadmap timeline, that's information — widen the MDE, expand the audience, or pick a different metric. Don't just run it anyway and hope the dashboard turns green early.
See more like this
ABWatcher catches A/B tests like this every day.
Watch live experiments at 1,000+ high-converting brands, complete with hypothesis and takeaway. Free forever for ten watched companies.