How do I know if my A/B test result is statistically significant?

Compare the two conversion rates with a two-proportion z-test. Divide the gap by its standard error to get z, then turn z into a p-value, the chance of a gap that big from luck alone. Below 0.05 is the usual bar. A 5.0% vs 5.6% test needs about 20,000 visitors per version to clear it.

A2% (± 0.27)B2.2% (± 0.29)Not significant yet: p = 0.32. Luck alone makes a gap this big 32% of the time.z = 0.986, lift +10%
Conversion rate A
2%
Conversion rate B
2.2%
Lift (B vs A)
10%
p-value (two-sided)
×0.324
z-score
0.986

Challenge: Make B significantly better using no more than 5,000 visitors in total, with a lift of 50% or less.

Sign-ups, purchases, clicks: whatever counts as a win.

Play

Load the new-button test, then double the traffic with the same rates.

Challenge: Make B significantly better using no more than 5,000 visitors in total, with a lift of 50% or less. The box under the picture turns green when you get it.

Stuck? Pick one of the examples from the “Try an example” menu, or press “New example.”

Understand

z=pB−pAp^(1−p^)(1nA+1nB)z = \frac{p_B - p_A}{\sqrt{\hat p(1 - \hat p)\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}}

Even two identical pages won't convert at exactly the same rate. The question is whether the gap is bigger than luck would make it. The standard error measures typical luck for this much traffic, and z is the gap measured in standard errors:

z=pB−pAp^(1−p^)(1nA+1nB)z = \frac{p_B - p_A}{\sqrt{\hat p(1 - \hat p)\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}}

The p-value is the chance of a gap at least that big from luck alone. The bars show each rate with its 95% range.

Use

Every input has a unit menu, so you can type values in the units you already have. Results follow your units.

Show the work

  1. Each version’s ratep_A = 2\%,\quad p_B = 2.2\%
  2. Pooled rate (if A and B were really the same)\hat p = \frac{200 + 220}{10{,}000 + 10{,}000} = 2.1\%
  3. Standard error of the gap\text{SE} = \sqrt{\hat p(1 - \hat p)\left(\tfrac{1}{n_A} + \tfrac{1}{n_B}\right)} = 0.2028\ \text{points}
  4. z-score: the gap measured in standard errorsz = \frac{0.2}{0.2028} = 0.9863
  5. Two-sided p-value: the chance of a gap at least this big by luckp = 2\,(1 - \Phi(|z|)) = 0.324

Export

Export PDF

Type in the visitors and conversions for each version.

  • This is the standard two-sided test for one A/B comparison, with the usual 5% threshold.

For learning and estimation. Verify with applicable codes, standards, and a qualified professional before using in design, construction, or safety-critical work.

Cheat card

p^=cA+cBnA+nB\hat p = \frac{c_A + c_B}{n_A + n_B}
z=pB−pAp^(1−p^)(1nA+1nB)z = \frac{p_B - p_A}{\sqrt{\hat p(1 - \hat p)\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}}
p-value=2 (1−Φ(∣z∣))p\text{-value} = 2\,(1 - \Phi(|z|))
SymbolMeaningUnit
nA,nBn_A, n_Bvisitors to each version
cA,cBc_A, c_Bconversions on each version
zzthe gap in standard errors
  • A z-score above 1.96 (or below −1.96) means p is below 0.05.
  • Small lifts need big samples. Halving the lift you want to detect takes about four times the traffic.
  • Decide the sample size before you start, and don't stop early when it looks good.

Print this card

Where it’s used

  • Computer Science
    Product teams test page designs, prices, and emails this way before rolling them out.
  • Finance & Business
    Marketers test ad copy and offers to find which one really converts better.
  • Money & Shopping
    Run a fair test on your own shop or newsletter before trusting a small bump.

Questions people ask

What is a p-value?

The chance of seeing a difference at least this big if A and B were really the same. A small p-value means luck is an unlikely explanation.

What does statistically significant mean?

The p-value is below a threshold set in advance, usually 0.05. It doesn't mean the difference is big or worth the change. That's a separate business question.

If the 95% ranges overlap, is the result not significant?

Not necessarily. Two ranges can overlap a little while the difference is still significant. The z-test on the difference is the right check.