Did your variant really win, or is it just noise? Enter the visitors and conversions for each version and this calculator returns the conversion rates, the relative uplift, and whether the difference is statistically significant — using a two-proportion z-test.
What this free tool is great for: a quick, one-off job with no signup — it runs entirely in your browser, so nothing leaves your device and there's nothing to manage.
Its honest limit: it's a one-off calculation in your browser — it doesn't save your scenarios, update as your real numbers change, or connect to your live accounts, so you re-enter the figures every time and can't watch how they move.
Statistical significance — a p-value below 0.05, i.e. 95% confidence — means the difference between your variants is unlikely to be random noise. It's a guard against fooling yourself: without it, you'll crown winners that are just the coin landing heads a few extra times. But significance is necessary, not sufficient — it says the difference is probably real, not that it's big enough to care about.
Hitting 95% does not mean the test is done. You still need enough sample size and a full business cycle — often two weeks — so weekday-versus-weekend and payday patterns wash out. Early significance on tiny numbers is the classic trap: results swing wildly at low sample sizes and only settle as data accumulates. If it isn't significant yet, keep it running.
Peeking and stopping early: checking daily and stopping the moment you touch 95% badly inflates your false-positive rate — look often enough and the line gets crossed by chance. Testing too many things at once: run twenty tests and one will look significant on pure luck. Ignoring practical significance: a statistically real 0.2% lift may not be worth the work to ship. Decide your sample size and duration before you start, then hold your nerve.
Pick one clear metric and one change. Estimate the sample size you need up front — big effects need little data, small ones need a lot. Run full weeks, not partial ones. Don't stop early because it looks good, and don't run forever hunting for significance that isn't there. A test that ends inconclusive is a real result too: it usually means the change simply didn't matter.
Tests fail in the design stage more often than in the maths. The discipline that keeps results interpretable: one primary metric declared before launch (the thing you'll actually decide on), one clear hypothesis ("moving the form above the fold will raise signups because fewer people scroll than we assume"), and ideally one change per variant. Test five differences at once and a winning variant teaches you nothing about *why* — you've bought a result but no understanding, and the next test starts from zero again. Bundled "redesign vs old" tests are sometimes justified; just know you're trading learning for speed, and label the result accordingly.
A p-value below 0.05 says the difference is probably real — it says nothing about whether it matters. With enough traffic, a 0.1% lift reaches perfect statistical significance while being worth less than the engineering time to ship it. The reverse trap exists too: an inconclusive test on low traffic doesn't prove "no effect", only "no effect big enough to detect with this sample". Before every test, decide the minimum lift worth acting on; after every test, compare the measured effect (with its uncertainty range) against that bar. The p-value is a gatekeeper against self-deception, not a measure of value.
Veteran testers know the pattern: a variant wins with +18%, ships to everyone, and the dashboard shows +7%. This isn't fraud — it's statistics. Tests that stop at the moment of significance systematically overestimate effects (the winner's curse: you declared victory on a lucky high), and novelty effects fade as returning visitors get used to the change. Protect yourself by running full pre-planned durations rather than stopping on a good day, and treat the post-launch measurement as part of the test. Budget mentally for shipped effects to be smaller than tested ones, and you'll stop being disappointed by real results.
The compounding asset of a testing program isn't any single win — it's the institutional memory of what was tried, what moved, and what did nothing. Without a log, teams re-test the button colour every eighteen months as personnel rotate, and the genuine insights (pricing-page layout matters, testimonials don't) evaporate. The log can be a simple sheet: date, hypothesis, variant screenshots, sample size, result, decision. Over a year it becomes the most honest document in the company about what your customers actually respond to — worth more than any individual 12% lift, because it steers every future bet.
Before believing any result, check the most boring number on the page: did the traffic split land where you set it? A 50/50 test that shows 54/46 in actual visitors has a sample-ratio mismatch — usually caused by a redirect that drops slow connections, a bot filter hitting one variant, or a caching layer serving one version unevenly. It sounds pedantic; it isn't. SRM means the groups weren't comparable, and every conclusion downstream inherits the bias — often in the exact direction that flatters the new variant. The check costs ten seconds (the calculator shows your visitor counts side by side), and among professional testers it's the first sanity gate every result must pass.
This calculator scores a finished test; the work is everything around it — building variants without engineering time, splitting traffic cleanly, tracking goals, watching for sample-ratio bugs, and keeping several experiments from contaminating each other. That's where VWO does more: a visual editor for variants, automatic traffic allocation and significance tracking, plus the heatmaps and recordings that generate better hypotheses than guessing. Use this page to understand and verify results; use a platform when testing becomes a rhythm rather than an occasional event — because the compounding returns come from the rhythm. A steady cadence of one honest test at a time, logged and learned from, beats a heroic quarterly redesign every time it's been measured — and it's a lot easier to defend in the next budget round. And keep perspective on what testing is for: it is not a machine for squeezing decimal points, it is an insurance policy against confidently shipping things that make everything worse — which, in mature testing programs, turns out to be roughly a third of all well-intentioned ideas. The tests that 'fail' by winning nothing save you from exactly those, and that avoided damage never shows up in the win-log but pays for the whole practice.
It's the probability that the difference between your variants is real rather than random noise. A p-value below 0.05 means roughly 95% confidence the result isn't chance.
It depends on your baseline rate and the uplift you want to detect, but smaller effects need far more traffic. As a rule of thumb, don't call a test before a few hundred conversions per variant and a full business cycle.
Better not. 'Peeking' and stopping at the first significant moment inflates false positives. Decide your sample size up front and let the test run its course.
A two-proportion z-test (two-tailed), the standard approach for comparing two conversion rates. For multivariate tests or sequential analysis, use a dedicated platform like VWO.
Blogger, teacher or toolmaker? Put this calculator on your own page — free forever, no strings. Copy the snippet below (the credit link is appreciated and keeps the tool free):
This tool is free and runs entirely in your browser. The link above is an affiliate link: we may earn a commission if you sign up, at no extra cost to you, and it never changes our honest take.
New dossiers, cost-traps we found, and tools that earned a keep — no hype, no sponsored-disguised-as-advice. Unsubscribe anytime.