A/B test sample size calculator
Find out how many visitors you need before your results are worth acting on.
How to use the calculator to find your required sample size
These three choices affect your required sample size and, by extension, how long your test will need to run.
1. Which mode?
- Test Planning: Use this before launch. Enter your local baseline and business-relevant MDE to calculate visitors per variation and estimated duration.
- Test Analysis: Use this after data arrives. Enter observed results to evaluate the effect, uncertainty, statistical evidence, and always-on SRM check.
2. Which metric?
- Conversion rate: Use for binary outcomes such as signups, purchases, or clicks. Lower variance usually means a smaller required sample than revenue metrics.
- Revenue per visitor: Use when revenue impact matters more than CVR alone. Upload visitor-level values so the plan reflects observed variance.
- Products per visitor: Use for ecommerce tests where items purchased per exposed visitor is the primary outcome.
3. Which method?
- Frequentist: Commit to the fixed sample before launch and do not stop at the first significant read.
- Sequential: Plan against the same target MDE, then use always-valid evidence when you need to monitor a live conversion, revenue, or products test.
- Bayesian: Use the fixed-horizon sample as a stability reference, then evaluate conversion probability, expected loss, and the credible interval.
Use the exact page or flow baseline, and set MDE from business value before looking at the resulting sample.
How to read the required sample
The output is a traffic plan, not a forecast that your variant will win.
- Visitors per variation: This is the sample each group needs before the test can reliably detect your target MDE.
- Duration: Duration converts the required sample into calendar time using daily traffic and allocation.
- MDE curve: The curve shows which effects become detectable as sample accumulates.
Do not game the MDE
Raising MDE only to fit a calendar date makes the test less sensitive.
- Use a business-relevant MDE before looking at sample size.
- If the sample is too large, revisit traffic allocation or test scope before changing the statistical target.
- Always keep SRM checks enabled when the test moves into analysis.
Sample size methodology
The calculator sizes your test from baseline, target MDE, confidence, power, tails, variants, and correction method.
- Conversion-rate planning: For conversion-rate metrics, the calculator uses two-sample proportion planning aligned to the backend power engine. Use this for signup, purchase, click, or activation tests.
- Continuous metrics: For revenue and products, upload baseline CSV data so the sample size reflects observed variance instead of a generic assumption. Use this when RPV, AOV guardrails, or products per visitor determine the decision.
- Multiple variants: Bonferroni and Sidak corrections adjust the threshold when you compare more than one variant. Use this for A/B/n tests where more comparisons increase false-positive risk.
Inputs that move the sample
These are the levers that make tests shorter, longer, more sensitive, or more conservative.
- Baseline rate: The current value of the metric you are testing. Local baselines beat site-wide averages.
- MDE: The minimum effect worth detecting. It is an input that represents business sensitivity.
- Power: The chance of detecting the target effect if it is real. Higher power means a larger sample.
Related calculators
Common questions about sample size
Use these answers to avoid underpowered tests and artificial launch deadlines.
- How do I choose the right MDE? Choose the smallest lift that would change a business decision. MDE is not the lift you expect or desire. A 5-10% relative effect can be a starting point only when you do not yet have a metric-specific business threshold.
- Why is my required sample size so high? A small MDE is usually the main driver because halving MDE roughly quadruples sample size. If the target is still worth detecting, run longer, increase qualifying traffic or allocation, or narrow the test to a higher-traffic flow.
- What is statistical power and why does it matter? Power is the probability of detecting your target effect when it is real. Higher power reduces false negatives but requires more visitors. Choose it before launch and keep it fixed.
- How does the number of variants affect sample size? Each additional variant adds traffic and comparisons. Bonferroni or Sidak correction makes the evidence threshold more conservative, so focused A/B tests are usually more sample-efficient than broad A/B/n tests.
- Which inputs have the biggest effect on sample size? MDE is usually the strongest lever, followed by baseline performance and metric variance. Power, confidence, test direction, allocation, variants, and multiple-comparison correction also change the requirement.
- What is statistical significance? Statistical significance means the observed result crossed your prespecified error threshold under the null hypothesis. It does not tell you whether the effect is large enough to matter.
- Should I use a one-tailed or two-tailed test? Use two-tailed by default so regressions and improvements are both detectable. A prespecified one-tailed test can require less sample for the target direction at the same alpha and power, but it gives up a formal claim in the opposite direction.
- What is Sample Ratio Mismatch? SRM checks whether observed visitors match the planned allocation. If it is flagged, investigate randomization, targeting, consent, bots, and instrumentation before trusting the result.
- What are Type I and Type II errors? A Type I error is a false positive and is controlled by the significance threshold. A Type II error is a false negative and is controlled by statistical power at the target MDE. Reducing either error generally requires more sample.