A/B test duration calculator

Find out how long to run your test before the results are worth trusting.

How to use the calculator to determine test duration

Metric variance and statistical method affect the required sample. Daily qualifying traffic and allocation then translate that sample into calendar time.

1. Which metric?

  • Conversion rate: Use for binary outcomes such as signups, purchases, or clicks. Lower variance usually means a shorter test than a revenue metric.
  • Revenue per visitor: Use when revenue impact matters more than CVR alone. Upload visitor-level values because higher variance generally lengthens the test.
  • Products per visitor: Use when items purchased per exposed visitor is the decision metric. Upload baseline values so their variance is reflected.

2. Which method?

  • Frequentist: The duration is the fixed-horizon estimate. Commit to the sample before launch and do not stop at the first significant read.
  • Sequential: The same MDE-based duration becomes a maximum planning reference while always-valid analysis may support an earlier stop.
  • Bayesian: The fixed-horizon duration is a stability reference; the eventual conversion decision uses probability, expected loss, and uncertainty.

3. Traffic inputs

  • Daily qualifying traffic: Use visitors to the tested page, audience, or flow, not total site sessions. Match the period used for the baseline.
  • Traffic allocation: Only visitors included in the experiment count. At stable traffic, halving allocation roughly doubles duration.

Use duration as a planning constraint, not as a reason to weaken the statistical target.

What your duration estimate means

Test duration is an estimate with assumptions behind it. Check the sample, traffic source, and sensitivity before building a roadmap around the date.

  • Test duration: The estimated calendar time to reach the required sample at the entered traffic and allocation. Fixed-horizon tests should not stop at first significance.
  • Required sample: The per-variation sample is the statistical requirement. Duration is that sample divided by effective daily traffic for each variation.
  • MDE curve: The curve shows which effect sizes become detectable as sample accumulates. Use it to understand the sensitivity traded away by a shorter run.

Before you act on the estimate

A deadline is not a statistical threshold, and the traffic assumptions can move.

  • Use traffic for the exact tested flow; site-wide traffic makes the duration look artificially short.
  • Recalculate when promotions, seasonality, eligibility, or allocation materially change daily traffic.
  • Cover complete business cycles relevant to your audience even when the sample is reached mid-cycle; this is an operational guardrail, not a universal fixed-week rule.

How test duration is calculated

The calculator derives sample size from baseline, MDE, confidence, power, variants, and correction, then divides by effective daily traffic per variation.

  • Frequentist duration: The duration estimates when the prespecified fixed-horizon sample should be reached at the current qualifying traffic and allocation. Use this when the team can commit to one planned analysis horizon.
  • Sequential duration: The same MDE-based duration anchors the maximum traffic commitment while always-valid evidence is monitored in analysis. Use this when an earlier conclusive read has operational value.
  • Bayesian duration: The fixed-horizon duration is retained as a practical data-stability reference for the Bayesian conversion decision. Use this before monitoring probability, expected loss, and the credible interval.

The inputs that decide run time

Duration changes when traffic, allocation, MDE, or metric variance changes.

  • Daily traffic: Qualifying visitors per day for the tested experience.
  • Traffic allocation: The percentage of qualifying traffic included in the experiment.
  • End date: A traffic-based estimate, not a guarantee. Recalculate when traffic changes.

Related calculators

Common questions about A/B test duration

Run time is a consequence of sample size, traffic, allocation, and sensitivity.

  • How long should I run an A/B test? Run long enough to reach the required sample for your baseline, business-relevant MDE, confidence, power, and variants. Also cover the complete business cycles relevant to the audience so a temporary day-of-week or campaign mix does not dominate the result.
  • What happens if I stop early? In a fixed-horizon test, stopping before the planned sample lowers power and makes a first threshold crossing prone to regression. Configure the test as sequential before launch if valid early monitoring is required.
  • My test would take months. What should I do? Confirm the MDE is the smallest effect worth acting on, then consider higher allocation, a higher-traffic surface, broader qualifying traffic, or a more valuable hypothesis. Do not lower confidence after seeing the estimate.
  • Do the days of the week matter? Yes. Traffic and behavior can vary across weekdays, weekends, billing cycles, promotions, and holidays. Include complete cycles relevant to the audience and avoid treating a partial cycle as representative.
  • How does traffic allocation affect duration? Only allocated qualifying visitors contribute to the sample. With stable traffic and even assignment, moving from 50% to 100% allocation roughly halves the calendar time, provided QA, consent, and operational risk allow it.
  • Which inputs have the biggest effect on duration? MDE is usually the largest statistical lever because halving it roughly quadruples required sample. Daily traffic and allocation then convert sample into time; metric variance, baseline, power, confidence, and variants also matter.
  • What is Sample Ratio Mismatch? SRM checks whether observed traffic matches the planned allocation. It runs during analysis, and a flagged or skipped check blocks a trustworthy launch decision even when the calendar estimate was met.
  • What are Type I and Type II errors? A Type I error is a false positive and is controlled by the significance threshold. A Type II error is a false negative and is controlled by power at the target MDE. Reducing either error generally increases sample and duration.
Start your 15-day free trial now.
  • No credit card needed
  • Access to premium features
You can always change your preferences later.
You're Almost Done.
What Job(s) Do You Do at Work? * (Choose Up to 2 Options):
Convert is committed to protecting your privacy.

Important. Please Read.

  • Check your inbox for the password to Convert’s trial account.
  • Log in using the link provided in that email.
  • To ensure you receive your 30-day trial from our ambassador, please use the same browser to claim your account.

This sign up flow is built for maximum security. You’re worth it!