How do I know if my A/B test result is statistically significant?
Compare the two conversion rates with a two-proportion z-test. Divide the gap by its standard error to get z, then turn z into a p-value, the chance of a gap that big from luck alone. Below 0.05 is the usual bar. A 5.0% vs 5.6% test needs about 20,000 visitors per version to clear it.
- Conversion rate A
- 2%
- Conversion rate B
- 2.2%
- Lift (B vs A)
- 10%
- p-value (two-sided)
- ×0.324
- z-score
- 0.986
Challenge: Make B significantly better using no more than 5,000 visitors in total, with a lift of 50% or less.
Play
Load the new-button test, then double the traffic with the same rates.
Challenge: Make B significantly better using no more than 5,000 visitors in total, with a lift of 50% or less. The box under the picture turns green when you get it.
Stuck? Pick one of the examples from the “Try an example” menu, or press “New example.”
Understand
Even two identical pages won't convert at exactly the same rate. The question is whether the gap is bigger than luck would make it. The standard error measures typical luck for this much traffic, and z is the gap measured in standard errors:
The p-value is the chance of a gap at least that big from luck alone. The bars show each rate with its 95% range.
Use
Every input has a unit menu, so you can type values in the units you already have. Results follow your units.
Show the work
- Each version’s rate
p_A = 2\%,\quad p_B = 2.2\% - Pooled rate (if A and B were really the same)
\hat p = \frac{200 + 220}{10{,}000 + 10{,}000} = 2.1\% - Standard error of the gap
\text{SE} = \sqrt{\hat p(1 - \hat p)\left(\tfrac{1}{n_A} + \tfrac{1}{n_B}\right)} = 0.2028\ \text{points} - z-score: the gap measured in standard errors
z = \frac{0.2}{0.2028} = 0.9863 - Two-sided p-value: the chance of a gap at least this big by luck
p = 2\,(1 - \Phi(|z|)) = 0.324
Export
Type in the visitors and conversions for each version.
- This is the standard two-sided test for one A/B comparison, with the usual 5% threshold.
For learning and estimation. Verify with applicable codes, standards, and a qualified professional before using in design, construction, or safety-critical work.
Cheat card
| Symbol | Meaning | Unit |
|---|---|---|
| visitors to each version | ||
| conversions on each version | ||
| the gap in standard errors |
- A z-score above 1.96 (or below −1.96) means p is below 0.05.
- Small lifts need big samples. Halving the lift you want to detect takes about four times the traffic.
- Decide the sample size before you start, and don't stop early when it looks good.
Where it’s used
- Computer Science
Product teams test page designs, prices, and emails this way before rolling them out. - Finance & Business
Marketers test ad copy and offers to find which one really converts better. - Money & Shopping
Run a fair test on your own shop or newsletter before trusting a small bump.
Related exhibits
- Start here
Poll margin of error: why about 1,000 people gives ±3%
The margin shrinks with the square root of the sample. Four times the people only halves it.
- Start here
Percent change: up 50%, then down 50%, and you’ve lost money
A percent is always “of something.” Change the something and the same percent means a different amount.
- Also try
Dice odds: what are my chances of rolling it?
Why 7 shows up so often with two dice, and what that means for your next move.
Questions people ask
What is a p-value?
The chance of seeing a difference at least this big if A and B were really the same. A small p-value means luck is an unlikely explanation.
What does statistically significant mean?
The p-value is below a threshold set in advance, usually 0.05. It doesn't mean the difference is big or worth the change. That's a separate business question.
If the 95% ranges overlap, is the result not significant?
Not necessarily. Two ranges can overlap a little while the difference is still significant. The z-test on the difference is the right check.