Two questions, one calculator
Every ad test asks two things, in this order. Before it starts: how many visitors does it need to find the difference I care about? After it ends: is the gap between A and B real, or could it be chance? Use the two tabs above in that order. Working out the sample size first is what makes the answer to the second question mean something.
“Visitors” can be landing page visitors, link clicks or anyone else who got a fair chance to convert, as long as both versions are counted the same way. “Conversions” are the people who did the thing you are testing for.
How the significance test works
The calculator runs a two-proportion z-test, the large-sample test for comparing two rates described in the NIST/SEMATECH e-Handbook of Statistical Methods:
- Pooled rate p = (conversions A + conversions B) ÷ (visitors A + visitors B).
- z = (rate B − rate A) ÷ √(p × (1 − p) × (1/visitors A + 1/visitors B)).
- Two-sided p-value = 2 × (1 − Φ(|z|)), where Φ is the standard normal distribution.
- The interval for the difference is (rate B − rate A) ± z × √(rate A × (1 − rate A) ÷ visitors A + rate B × (1 − rate B) ÷ visitors B). The lift range divides that by rate A, which treats rate A as known, so it is an approximation.
Worked example. Version A gets 200 conversions from 10,000 visitors (2.00%). Version B gets 245 from 10,000 (2.45%), a 22.5% relative lift. The pooled rate is 445 ÷ 20,000 = 2.225%, the standard error is 0.209 percentage points, and z = 0.45 ÷ 0.209 = 2.157. The two-sided p-value is 0.031, below 0.05, so at 95% confidence B wins. The 95% interval for the difference runs from +0.04 to +0.86 percentage points: B is better, but the true lift could be anywhere from about 2% to 43%.
The test is an approximation that needs enough data. NIST’s handbook gives one criterion: at least 5 successes and 5 failures. The calculator holds back its verdict and says not enough data yet until each version has 5 conversions and 5 non-conversions. With fewer than that, NIST points to an exact test for small samples, but for an ad test the practical answer is to keep it running.
How the sample size is worked out
For a two-sided test, visitors per version:
n = (z1−α/2 × √(2 × p̄ × (1 − p̄)) + z1−β × √(p1(1 − p1) + p2(1 − p2)))² ÷ (p2 − p1)²
Here p1 is your current rate, p2 is that rate with the lift applied, and p̄ is their average. It is the normal approximation, without a continuity correction, and it has the same shape as the single-proportion sample size formula in the NIST handbook, with a second group added. At 95% confidence z1−α/2 is 1.959964, and at 80% power z1−β is 0.841621.
Worked example. A 2% conversion rate, and you want to catch a 20% lift, so 2.0% against 2.4%. The average rate is 2.2%. The top line is 1.959964 × 0.2074 + 0.841621 × 0.2074 = 0.5811, squared is 0.3377, and dividing by 0.004² gives 21,109 visitors per version, 42,218 in all. At 3,000 visitors a day that is 15 days.
The lift you ask for matters most. On the same 2% rate, a 30% lift needs 9,798 per version and a 10% lift needs 80,682. Halving the lift roughly quadruples the traffic, because the gap sits squared at the bottom of the formula.
Why stopping early inflates false positives
A 95% confidence level means that when A and B truly convert the same, a single test at a fixed sample size still calls a winner about 5% of the time. That 5% is only true if you look once, at the end.
Checking the result every day and stopping the first time it shows significant is a different procedure. The rates wander as data comes in, and each look is another chance for a chance gap to cross the line. Stop at the first crossing and you have picked the luckiest moment, so the real false-positive rate ends up above 5%, and it grows the more often you look. The p-value on the screen does not know you peeked, so it still reads as if you had looked once.
The fix is dull and it works: decide the sample size before the test starts, let it reach that size, and read the result once. Looking in the middle to check nothing is broken is fine. Stopping because the number looks good is not.
What Meta and Google Ads say about test length
- Meta’s A/B testing best practices recommend tests of at least 7 days, say tests shorter than that may produce inconclusive results, and cap A/B tests at 30 days. The same page suggests running longer if customers usually take more than 7 days to convert.
- Meta typically suggests an estimated power of 80% or higher, the default power here. Its A/B results show a confidence percentage, and it treats 65% or higher as a winning result for an A/B test. That figure comes from Meta’s own method and is not the same thing as the confidence level on this page, so the two will not match.
- Meta also advises against testing informally by switching ad sets on and off, because audiences can overlap, and recommends the same budget for both versions.
- Google Ads recommends letting an experiment run for 2 to 3 weeks, shows an 80% confidence interval by default with other levels on request, and marks significant results with a blue asterisk. Its statistical method uses jackknife resampling over 20 buckets per arm, so its intervals will not match a z-test exactly either.
What this calculator does not do
- It compares two versions. With three or more, each extra comparison is another chance of a false winner, and the simple test here does not adjust for that.
- It tests conversion rates, not revenue per visitor or order value, which need a different test.
- It is not a sequential test. It assumes you read the result once, at the sample size you planned.
- It cannot tell whether your split was fair. If one version got a different audience, budget or time of day, no formula fixes that.
If you would rather not run the numbers by hand
Hermoso is marketing on autopilot. Ask it what to scale, pause and test next and it reads every ad platform you have connected, puts campaigns with too little data in their own bucket instead of judging them, and gives each suggested test a hypothesis, the one thing to change, the metric that decides it, a budget and a duration. It also makes on-brand image and video variants to test, and builds campaigns on your own ad accounts, created paused. It works in the web app, or inside Claude, ChatGPT and Cursor through its MCP server. Once a test has a winner, the free ROAS calculator tells you whether it is also profitable.
Free tools: · ROAS calculator · Ad copy length checker
Frequently asked questions
How do I know if my ad A/B test is statistically significant?
Enter visitors and conversions for both versions. The calculator runs a two-proportion z-test and gives a two-sided p-value. At 95% confidence, a p-value below 0.05 means the gap is unlikely to be chance. For 200 of 10,000 against 245 of 10,000, p is 0.031, so B wins.
How many visitors does an A/B test need?
It depends on your conversion rate and the smallest lift you want to detect. At a 2% conversion rate, finding a 20% lift at 95% confidence and 80% power takes 21,109 visitors per version. A 10% lift takes 80,682.
Why does the calculator say not enough data yet?
The z-test is a large-sample approximation. The calculator waits until each version has at least 5 conversions and 5 non-conversions, one criterion given in the NIST e-Handbook, before it gives a verdict.
Can I stop a test as soon as it looks significant?
Not if you want the confidence level to mean what it says. Checking repeatedly and stopping at the first significant reading gives chance more opportunities to cross the line, so the real false-positive rate ends up higher than the one you chose. Set the sample size first and read the result once.
How long should a Meta A/B test run?
Meta recommends at least 7 days, says shorter tests may be inconclusive, and allows up to 30 days. Google Ads recommends 2 to 3 weeks for experiments. Use the sample size tab with your daily traffic to see how long your own test needs.
Is anything I type sent anywhere?
No. The sums run in your browser. There is no network request carrying what you type and nothing is stored.
Hermoso tells you what to scale, pause and test next across your ad accounts. Start free, 250+ earnable free credits, no card.
Start free → See pricing