Before you launch a test, find out how much traffic it actually needs. Enter your baseline conversion rate and the smallest uplift worth detecting, and this returns the visitors required per variant — so you don't call a test too early.
What this free tool is great for: a quick, one-off job with no signup — it runs entirely in your browser, so nothing leaves your device and there's nothing to manage.
Its honest limit: it's a one-off calculation in your browser — it doesn't save your scenarios, update as your real numbers change, or connect to your live accounts, so you re-enter the figures every time and can't watch how they move.
The single most common way to wreck an A/B test is to stop it too early. You launch a test, one version jumps ahead after a few hundred visitors, you declare victory and ship it — and the "win" quietly evaporates in production because it was never real, just an early run of luck. Calculating the sample size you need before you start is the cure. It tells you how much traffic the test must gather before its result means anything, which turns testing from wishful pattern-spotting into something you can actually trust. This calculator gives you that number; the ideas below explain why it's so much larger than your gut expects.
Sample size isn't arbitrary — it's driven by a few inputs, and understanding them makes you a far better tester. Your baseline conversion rate (how the current version performs) sets the starting point. The minimum detectable effect (the smallest improvement you want to be able to catch) is the biggest lever. Your desired statistical significance (usually 95%, your tolerance for false positives) and statistical power (usually 80%, your tolerance for false negatives) fill in the rest. Change any of these and the required sample moves, often dramatically. The calculator does the maths, but knowing which dial you're turning — and what you're trading — is what separates a rigorous test from a hopeful one.
The minimum detectable effect is where people fool themselves. It's the smallest lift you want the test to be able to detect, and here's the brutal trade-off: the smaller the effect you want to catch, the enormously larger the sample you need. Detecting a big, obvious improvement takes relatively little data; detecting a tiny 1% lift can take months of traffic. So you have to be honest up front about how small an effect is actually worth chasing. Setting an unrealistically tiny detectable effect on modest traffic dooms the test to never reach significance, while being realistic about what's worth detecting keeps your tests finishable.
Everyone worries about false positives — calling a win that isn't real — but false negatives matter just as much. Statistical power, conventionally set at 80%, is your protection against them: it's the probability your test will detect a real effect that genuinely exists. Under-power a test and you can run a genuinely better variant, fail to reach significance, conclude "no difference," and throw away a real improvement. Power is why sample size has to account for both kinds of error. A test sized only to avoid false positives, with no thought to power, will miss real wins as often as it catches them — which is its own expensive form of failure.
The uncomfortable truth for smaller sites is that reliable A/B testing needs a lot of traffic, and there's no way around it. If you get a few hundred visitors a week, a properly-powered test to detect a modest lift might need to run for months — long enough that seasonality and other changes muddy the result. This isn't a reason to give up, but it is a reason to be strategic: test big, bold changes that produce large detectable effects rather than fiddly tweaks, prioritise your highest-traffic pages, and accept that on low volume, some questions simply can't be answered by A/B testing and are better decided by judgement and qualitative research.
The whole point of sizing a test is that it happens before you start. You decide the baseline, the effect worth detecting, and your confidence and power levels, and you get a target sample and a rough duration. Then you commit to running until you hit it. Working this out afterwards — or not at all — is how tests get stopped on a hunch. A test with a pre-declared sample size and end condition removes the temptation to quit the moment the numbers look good, which is exactly the temptation that produces false wins. Plan the finish line before you fire the starting gun.
Once a test is running, the danger becomes peeking — checking the results repeatedly and stopping the instant they cross the significance line. This badly inflates your false-positive rate, because if you look often enough, random noise will eventually touch 95% by pure chance even when there's no real difference. The discipline is to decide your sample size and duration in advance and hold your nerve until you reach them, resisting the urge to call it early on a good-looking day. Continuous peeking with early stopping is statistically equivalent to running many mini-tests and keeping whichever one happened to look best — which is no test at all.
Beyond raw sample size, timing matters. Run a test for full weekly cycles, not partial ones, so weekday-versus-weekend and payday patterns wash out rather than skewing the result. A test that hits its sample size mid-week should usually run to the end of the week anyway. Traffic behaves differently on a Tuesday morning than a Saturday night, and a test that captures only part of that rhythm can crown a winner that only wins under those specific conditions. Sizing the sample gets you enough data; running full cycles makes sure that data represents your real, varied traffic rather than a lucky slice of it.
Knowing the sample size is the planning half of testing — this calculator gives you that number and the discipline that comes with it. Actually running the experiment — building the variants without a developer, splitting traffic, tracking conversions and reaching significance across many tests — is a separate, ongoing job. That's where a platform like VWO does more: it lets you create and run A/B tests visually, manages the traffic split and the statistics, and turns testing into a repeatable programme rather than a one-off. Use this tool to size a test honestly before you start; use a testing platform to build and run it properly once you have your number.
Power (usually set at 80%) is the chance your test detects a real effect if one exists. Higher power needs more visitors but reduces the risk of missing a genuine winner.
It's the smallest improvement you care about catching, as a relative percentage. Detecting a 20% lift needs far less traffic than detecting a 5% lift — be realistic about what's worth testing.
Conversion data is noisy, so distinguishing a real difference from random chance takes volume — especially for small effects or low baseline rates. The calculator shows the honest number.
Use an experimentation platform like VWO to build variants, split your traffic and measure results — this calculator just tells you how big the test needs to be.
Blogger, teacher or toolmaker? Put this calculator on your own page — free forever, no strings. Copy the snippet below (the credit link is appreciated and keeps the tool free):
This tool is free and runs entirely in your browser. The link above is an affiliate link: we may earn a commission if you sign up, at no extra cost to you, and it never changes our honest take.
New dossiers, cost-traps we found, and tools that earned a keep — no hype, no sponsored-disguised-as-advice. Unsubscribe anytime.