The statistical significance calculator for A/B tests

Calculate sample size, statistical significance, and test duration in one place. Know exactly when to launch and when to keep testing.

Three decisions to make before you start

Before you enter any inputs into the stat sig calculator, decide which test mode fits your situation, which metric you are optimizing, and which statistical method your team works with.

1. Which mode?

  • Test Planning: You have not launched yet. Enter your baseline rate and business-relevant MDE to find out how many visitors you need and how long to run.
  • Test Analysis: Your test is done or already running. Enter observed results to check p-values, confidence intervals, SRM, and method-specific decision signals.

2. Which metric?

  • Conversion rate: Best for binary outcomes like signups, purchases, or clicks. Converted/not-converted has lower variance than revenue.
  • Revenue per visitor: Use this when revenue impact matters more than CVR alone. Upload visitor-level revenue data so variance is estimated from your traffic.
  • Products per visitor: Use this for ecommerce tests where items purchased per visitor is the primary outcome.

3. Which method?

  • Frequentist: Set a fixed sample size upfront and do not stop at first significance. This is the standard approach for defensible launch decisions.
  • Sequential: Use always-valid monitoring when you need to check a live test as data arrives.
  • Bayesian: Use probability-to-win, expected loss, and credible intervals for conversion-rate decisions.

Full methodology detail for each method is in the Statistical methodology section below.

What your result means

Use these checks before you act on a sample size, duration, or significance result.

  • Required sample: Visitors needed per variant before your result is valid. Do not stop before you hit this number, even if you see significance early.
  • Test duration: Estimated days to reach the required sample at your current traffic and allocation. If it runs beyond six to eight weeks, revisit the MDE, scope, or allocation.
  • MDE curve: Shows how the minimum detectable effect shrinks as your sample grows. Use it as the output view of test sensitivity.

Before you act on the result

The most common ways teams misread calculator output.

  • Do not stop when you first see significance. Wait for the planned sample unless you are using the sequential monitoring path.
  • Run the SRM check first. If the traffic split is mismatched, the p-value is not reliable.
  • Make sure the primary metric matches the business decision. CVR wins can hide revenue regressions.

Statistical methodology, demystified

Three calculation paths. Pick the one that matches how your team actually decides to ship.

  • Frequentist: Conversion-rate analysis uses a two-sample proportion test. Revenue and products use Welch-style continuous-metric analysis from uploaded data. Bonferroni or Sidak corrections control multiple comparisons. Use this when you set a fixed sample size before launch and will not check early.
  • Sequential: Sequential analysis reports an always-valid p-value during post-test monitoring. Planning sample size remains anchored to the same target MDE baseline. Use this when you need to monitor a live test and stop only when the evidence is conclusive.
  • Bayesian: Bayesian conversion analysis uses Beta posteriors to estimate probability to win, expected loss, and credible intervals. Use this when probability language and risk make the decision clearer for stakeholders.

What each input and output means

Reference these while you are running your calculations.

  • Sample size and MDE: MDE is not the lift you expect or hope to see. Start with the smallest lift that would actually matter for the business.
  • P-values and confidence: A p-value is the probability, assuming no true effect, of observing a result at least as extreme as this one. Pair it with the effect interval and your business-relevant MDE.
  • Power and error control: At 80% power, a test has an 80% chance of detecting the target MDE when that effect is real under the model assumptions. Higher power reduces false negatives but needs more visitors.
  • Sample Ratio Mismatch: SRM checks whether observed visitor counts match the planned allocation. If SRM is flagged, investigate before trusting the result.
  • Conversion rate: Binary outcomes have lower variance than revenue metrics, so conversion-rate tests usually reach a decision faster.
  • RPV, AOV, and products: Revenue Per Visitor keeps the visitor as the unit of analysis. AOV is useful as a guardrail when order size changes matter.

Related calculators

Common questions about A/B test statistics

Starting your test? These are the questions most teams get wrong before they run the numbers.

  • What is MDE and how do I choose the right number? MDE, or minimum detectable effect, is the smallest improvement you want the test to be able to catch. It is not your expected lift or desired lift. Pick an MDE that reflects the smallest improvement actually worth shipping.
  • When should I use Bayesian vs. Frequentist vs. Sequential? Use Frequentist for fixed-horizon tests, Sequential when you need always-valid monitoring, and Bayesian when probability-to-win and expected loss make the decision clearer.
  • How do I calculate sample size for a conversion-rate test? Sample size depends on baseline conversion rate, minimum detectable effect, statistical power, confidence level, test direction, and the number of variants.
  • What is statistical significance? Statistical significance means the observed result crossed your prespecified error threshold under the null hypothesis. It does not tell you the probability that the variant is better or whether the effect matters to the business.
  • Should I use a one-tailed or two-tailed test? Use a two-tailed test by default because it detects both improvements and regressions. Use one-tailed only when the direction was prespecified and the opposite direction would genuinely not change the decision.
  • What is Sample Ratio Mismatch? SRM checks whether visitor counts match the planned allocation. A mismatch can point to randomization bugs, targeting mistakes, bot traffic, or analytics collection issues. If SRM is flagged, do not trust the result regardless of its p-value.
  • What are Type I and Type II errors? A Type I error can ship a false winner; a Type II error can discard a change that truly matters. The significance threshold limits false positives, while power at your target MDE limits false negatives. Choose both before launch so the decision rule reflects the cost of each mistake.
  • What is a good conversion rate for my website? There is no universal good conversion rate. Compare the same audience, intent, device mix, and funnel over time. Your own stable baseline is more useful for experiment planning than a broad industry benchmark.
Start your 15-day free trial now.
  • No credit card needed
  • Access to premium features
You can always change your preferences later.
You're Almost Done.
What Job(s) Do You Do at Work? * (Choose Up to 2 Options):
Convert is committed to protecting your privacy.

Important. Please Read.

  • Check your inbox for the password to Convert’s trial account.
  • Log in using the link provided in that email.
  • To ensure you receive your 30-day trial from our ambassador, please use the same browser to claim your account.

This sign up flow is built for maximum security. You’re worth it!