A/B Test Significance Calculator
So your a/b test variant is beating control. Before you publish it, it's worth checking whether that lift will hold up or whether it's just random swings in traffic. Plug your visitors and conversions into the calculator below to see your p-value, confidence level, and the likely range of your true lift. Haven't launched yet? Switch to Plan a test to find out how much traffic you'll need to get a result you can trust.
What Is an A/B Test Significance Calculator?
An A/B test significance calculator compares the conversion rates of two versions of a page, control (A) and variant (B). It tells you whether the difference between them is statistically meaningful or just normal traffic fluctuation. It needs four inputs: visitors and conversions for each version.
It returns the conversion rate for each variant and the lift, shown as both an absolute percentage-point difference and a relative percentage. It also returns the p-value, the confidence level, and a confidence interval for where the true difference likely falls. Used together, these tell you whether a result is reliable enough to publish.
Our calculator runs a two-proportion z-test, which is the standard method behind most A/B testing tools. The test assumes three things. Each visitor sees only one variant. Each variant has a reasonable number of conversions, usually 30 or more. And your traffic reflects normal site behavior. If those assumptions break, the output can't be trusted, whatever the p-value says. Examples include testing during a holiday spike or a contaminated traffic split.
How to Use This Calculator
Grade a test checks a test that is running or finished. Enter visitors and conversions for control and variant, and choose your confidence level. The calculator tells you whether the variant truly won, how confident you can be, and how big the lift is.
Plan a test is for before launch. Enter your baseline conversion rate, the smallest lift worth detecting, and your daily traffic. You'll get the sample size needed per variant and an estimate of how many days the test should run. Use it before every test. Most failed tests were set up to fail before they started.
How to Calculate Statistical Significance for an A/B Test
This is the math the calculator runs, with an example.
1. Calculate each conversion rate. Conversion rate = (conversions ÷ visitors) × 100. Suppose control gets 480 conversions from 10,000 visitors (4.8%). The variant gets 520 conversions from 10,000 visitors (5.2%). (New to the formula? Our conversion rate calculator covers it in depth.)
2. Find the lift, both ways. Absolute lift is the raw difference: 5.2% − 4.8% = +0.4 percentage points. Relative lift is that difference as a share of baseline: 0.4 ÷ 4.8 = +8.3%. Always look at both. We've seen teams celebrate an "18% improvement" that turned out to be +0.3 percentage points. Relative lift makes results sound bigger. Absolute lift shows what actually happens to revenue.
3. Calculate the pooled standard error. Combine both variants to get a pooled rate. Here that's 1,000 ÷ 20,000 = 5%. Then work out how much the two rates would normally vary if there were no real difference. In this example, the standard error is about 0.31 percentage points.
4. Get the z-score. Divide the difference by the standard error: 0.4 ÷ 0.31 ≈ 1.30. The rates are about 1.3 standard errors apart.
5. Convert to a p-value. A z-score of 1.30 gives a two-tailed p-value of about 0.19, well above the 0.05 threshold.
The verdict: not significant. An 8.3% relative lift across 20,000 visitors sounds convincing, but it's still inside normal variation. The 95% confidence interval runs from roughly −0.2 to +1.0 percentage points, and that range includes zero. Using the Plan tab, detecting this lift reliably would take about 46,600 visitors per variant. If the same rates held at 50,000 visitors per variant, the p-value would drop to about 0.004, a clear win.
A note on small samples: with fewer than 30 conversions per variant, the z-test's normal approximation stops being accurate and you need exact methods instead. Most online calculators don't handle this case. If you're that low, collect more data before trusting any result.
p-Value, Confidence Level, and Confidence Intervals Explained
p-value. The p-value is the probability of seeing a difference at least this large if the variants actually performed the same. A p-value of 0.03 means that if nothing had changed, a gap this big would appear only 3% of the time. That's strong evidence for the winner. A p-value of 0.12 means 12%, which is too high to credit the design change over ordinary fluctuation.
Confidence level. The standard is 95% (α = 0.05), but that's a convention, not a law. Low-traffic sites often use 90% (α = 0.10), because waiting for 95% can take months. The trade-off is a higher false-positive risk, about 10% instead of 5%. For high-stakes changes such as checkout redesigns or pricing tests, 99% (α = 0.01) is worth the extra traffic, because shipping a false winner there costs real revenue.
Confidence interval. The confidence interval is the range where the true lift likely falls. An interval of +0.2 to +0.8 percentage points means you can be 95% confident the real lift is somewhere in that range. Its width tells you something the p-value doesn't: precision. A significant result with an interval of +0.1 to +2.5 points is real, but you still don't know how big it is. That uncertainty should affect how much you invest in rolling it out.
A/B Test Sample Size: Plan Before You Launch
Stopping tests before they reach an adequate sample is the most common testing mistake we see. With small samples, p-values swing wildly. A test showing p = 0.04 on Tuesday can show p = 0.08 on Wednesday after a few hundred more visitors.
Required sample size depends on three things: your baseline conversion rate, your minimum detectable effect (the smallest lift worth catching), and your confidence level. Smaller effects need far more traffic. Halving the lift you want to detect roughly quadruples the visitors you need.
Test duration (days) = total required visitors ÷ daily visitors.
For a site converting at 3%, at 95% confidence and 80% power:
|
Lift you want to detect |
Visitors per variant |
At 10K visitors/week |
At 2K visitors/week |
|
+10% relative (3.0% → 3.3%) |
~53,200 |
~11 weeks |
~53 weeks |
|
+20% relative (3.0% → 3.6%) |
~13,900 |
~3 weeks |
~14 weeks |
|
+30% relative (3.0% → 3.9%) |
~6,500 |
~2 weeks* |
~7 weeks |
*Run every test for at least two full weeks, even if you hit your sample size sooner. That captures weekday and weekend behavior.
The takeaway is uncomfortable but useful. Unless you have serious traffic, small tweaks aren't testable. Lower-traffic stores should test bold changes that could plausibly move conversion 20% or more. Plug your own numbers into the Plan a test tab.
Multivariate tests multiply the problem. Three headlines × two CTAs × two images = 12 combinations. If a simple A/B test needs 4,000 total visitors, that multivariate test needs 24,000. Save multivariate testing for very high-traffic pages where you need to understand how elements interact.
Common A/B Testing Mistakes That Fake a Win
Peeking and stopping early. If you check results daily and stop the moment p drops below 0.05, you're catching random favorable swings, not real effects. This is called optional stopping. It can push your actual false-positive rate past 25%, even though you think you're working at 95% confidence. Set your sample size in advance and don't make a call until you reach it.
Sample ratio mismatch (SRM). If you set up a 50/50 split but ended up at 52/48 or worse, something is broken. Common causes are uneven bot filtering, a bug in the testing tool, or tracking code that fires inconsistently. Run a chi-square test on the split. If the p-value is below 0.01, you have SRM, and your significance result can't be trusted. Fix the tracking and restart with clean data.
Confusing statistical and practical significance. With enough traffic, even a tiny lift becomes statistically significant. We've seen tests hit p < 0.05 with a +0.08 percentage-point lift that worked out to two extra conversions a week. That isn't worth the QA time, never mind the development and maintenance. Before shipping any winner, ask: if this lift is real, does it justify the work? If it doesn't, move on to a bigger test.
Statistical Significance Tells You If a Test Worked, Not What to Test
A significance calculator protects you from shipping noise. It can't tell you whether you're testing the right things in the first place, and that's where most teams lose.
Most testing roadmaps are full of button colors, headline variations, and microcopy. These tests can reach significance and still produce lifts too small to matter. As the sample size table shows, small lifts also take months to detect on most sites. The real opportunity is finding the high-friction problems that are actually costing you conversions.
We've analyzed conversion funnels for over 10,000 brands. The issues that move the needle usually aren't subtle. They become obvious once you look at the site the way your customers do: navigation that hides critical product information, product pages that bury social proof or don't answer basic questions, checkout flows with unnecessary steps, and mobile layouts where key buttons are hard to tap.
Our Conversion Reports find those issues, rank them by likely impact, and deliver dev-ready Figma designs with the reasoning behind each recommendation. Every report includes A/B testing guidance with clear success metrics. We don't claim every recommendation will win. We do claim the changes are worth testing, because they come from behavioral patterns we've seen hurt conversion across thousands of sites. Brands like Braxley Bands have increased conversion by 40% this way.
Want to see what's killing your conversion rate? Try Oddit Free. We'll redesign one section of your site, with a full conversion report, at no cost. No credit card, no meetings.
A/B Testing FAQs
What is a good p-value for an A/B test?
Below 0.05 is the standard threshold for significance at 95% confidence. Lower-traffic sites sometimes accept 0.10 (90% confidence), and high-stakes tests like pricing or checkout should aim for 0.01 (99% confidence).
How many visitors do I need for an A/B test?
It depends on your baseline conversion rate and the smallest lift you want to detect. A site converting at 3% needs about 13,900 visitors per variant to detect a 20% relative lift at 95% confidence. Use the Plan a test tab above for your own numbers.
How long should an A/B test run?
Run it until you reach your planned sample size, and for at least two full weeks so you capture weekly traffic cycles. Stopping early because results "look good" is when false positives are most likely.
What's the difference between absolute and relative lift?
Absolute lift is the raw percentage-point difference (4.8% → 5.2% = +0.4 points). Relative lift is that change as a share of baseline (+8.3%). Absolute lift shows the real business impact.
Can a test be statistically significant but not worth shipping?
Yes. With enough traffic, even tiny lifts reach significance. If the lift is too small to justify the effort to build and maintain the change, skip it and test something bigger.