Sequential A/B testing calculator
Monitor your test as it runs. Stop early when the evidence is there.
Monitor a live test without fixed-horizon peeking
Sequential analysis produces an always-valid p-value for repeated reads. Planning stays anchored to the same business-relevant MDE so you enter the test with a clear traffic commitment.
- Always-valid evidence: The sequential p-value is designed to remain valid as the test is checked over time, unlike repeatedly applying a fixed-horizon test.
- A planning reference: The maximum sample shown here is the fixed-horizon sample for the selected MDE, confidence, power, and allocation. It anchors sensitivity; it is not an analysis cadence.
- A real data-quality gate: A threshold crossing is actionable only after the always-on SRM check has enough traffic and reports a clean allocation.
How to use the sequential testing calculator
Sequential testing changes how you read a live conversion test. Set the planning target and operational decision rule before data starts arriving.
1. Which mode?
- Test Planning: Enter the local baseline and smallest business-relevant MDE to set the maximum sample commitment and estimated duration.
- Test Analysis: Enter live or completed visitors and conversions to get the always-valid p-value, effect interval, power reference, and SRM status.
2. Which metric?
- Conversion rate: This sequential view is for binary outcomes such as signups, clicks, purchases, or activations.
3. Decision rule
- Prespecify the stop rule: Stop only when the sequential threshold is crossed, SRM is clean, and the observed effect is worth acting on. Do not invent a visitor-based check schedule.
Use the planning result as the maximum traffic commitment, then monitor on a consistent operational schedule with always-valid p-values in analysis.
How to read sequential output
Sequential output is about monitoring discipline, not a shortcut around data quality.
- Maximum sample size: The sample shown in planning is the fixed-horizon reference for the selected MDE, confidence, and power, used here as the maximum traffic commitment.
- Always-valid p-value: In analysis, the p-value remains valid under repeated looks when using this sequential method. A crossing still needs a clean SRM check and a meaningful effect.
- SRM check: A clean traffic split is still required before any sequential result is trustworthy.
Before you stop early
Sequential monitoring only works when the evidence is conclusive and the traffic split is clean.
- Set the MDE and stop rule before launch; do not rewrite them after seeing the result.
- Monitor on a consistent operational schedule tied to real decision points, not an arbitrary visitor-count cadence.
- Do not stop, ship, or declare a winner while SRM is flagged or has too little traffic to run.
Sequential monitoring methodology
The sequential path reports always-valid post-test evidence and an MDE-based maximum sample commitment.
- Always-valid p-values: The calculator reports an e-value-based sequential p-value that remains valid under repeated checking. Use this for live tests where waiting for a single final read is operationally expensive.
- Maximum sample commitment: Planning uses the fixed-horizon sample for the same baseline, MDE, confidence, power, and variants as a transparent maximum commitment. Use this to set a clear ceiling before launch.
- Data quality first: Sequential evidence is not trustworthy if traffic allocation is broken. Use the SRM card before any stop or ship decision.
What sequential output means
Sequential monitoring changes how you check the result, not what data quality requires.
- Always-valid p-value: A p-value designed to remain valid as you monitor the test over time.
- Maximum sample size: The per-variation traffic commitment used as the ceiling before monitoring begins.
- Stop condition: A stop decision should require conclusive evidence and a clean SRM check.
Related calculators
Common questions about sequential A/B testing
Sequential testing helps with monitoring, but it does not relax data quality or metric discipline.
- What is the difference between sequential and standard A/B testing? A standard fixed-horizon test assumes the decision is made at a prespecified sample. This sequential path uses an always-valid p-value designed for repeated monitoring, so checking the live test does not create the same false-positive inflation as naive peeking.
- Does sequential testing require more traffic? This calculator does not add a separate ceiling premium. Planning uses the same fixed-horizon MDE sample as a maximum commitment; the sequential benefit is valid repeated monitoring and a possible earlier threshold crossing.
- Who should use sequential testing? Use it when a team needs to monitor a live conversion test and may stop when evidence becomes conclusive. It is especially useful when waiting for one fixed final read has a real operational cost.
- How often should I check a sequential test? The inference does not require a prespecified visitor cadence. Use a consistent operational schedule tied to real decision points, and avoid reacting to every small fluctuation in the data.
- What if I reach the maximum sample without crossing the threshold? Do not extend the test or relax the threshold after the fact. Report that the test did not establish a detectable effect at the chosen sensitivity, then review the confidence interval and business context rather than declaring control a winner.
- Can I use this page for revenue? This page is locked to conversion-rate testing. Use the main calculator for revenue and product metrics.
- When can I call my sequential test? Call it only when the always-valid threshold has crossed, the always-on SRM check is clean, and the observed effect satisfies the business decision rule set before launch. If SRM is flagged or skipped for low traffic, the decision is blocked.