Back to blog
Engineering7 min read

Bayesian vs. Frequentist A/B Testing: A Practical Guide

M

Michał Pogoda-Rosikoń

Founder · 2026-04-05

Bayesian vs. Frequentist A/B Testing: A Practical Guide

Our first significance engine was a z-test. It worked, mostly.

When I built the first version of SplitMonk's stats engine, I did what every textbook tells you: two-proportion z-test, alpha = 0.05, compute a p-value, done. Standard frequentist approach. It took about 40 lines of TypeScript and I shipped it in an afternoon.

Then the complaints started.

"Why does it say 'not significant' when variant B clearly has more conversions?" (Because the sample size is too small.) "I checked yesterday and it was significant, now it's not -- what happened?" (Because you peeked and the result regressed to the mean.) "What does p=0.048 actually mean? Can I ship this?" (It means... well, let me write a blog post about it.)

The z-test was technically correct. But it was producing outputs that humans couldn't use properly. So we switched to Bayesian. Here's the whole journey.

The frequentist approach: what we started with

Here's roughly what our first computeSignificance() looked like:

function frequentistTest(
  controlConversions: number, controlVisitors: number,
  variantConversions: number, variantVisitors: number
) {
  const p1 = controlConversions / controlVisitors;
  const p2 = variantConversions / variantVisitors;
  const pPooled = (controlConversions + variantConversions)
    / (controlVisitors + variantVisitors);

  const se = Math.sqrt(pPooled * (1 - pPooled)
    * (1/controlVisitors + 1/variantVisitors));
  const z = (p2 - p1) / se;
  const pValue = 2 * (1 - normalCdf(Math.abs(z)));

  return {
    significant: pValue < 0.05,
    pValue,
    lift: (p2 - p1) / p1,
  };
}

This works fine if you follow the rules perfectly: decide your sample size upfront, run to completion, analyze once. In practice, nobody does this. Product managers peek. Founders peek. I peek. Everyone peeks.

And every time you peek, you inflate your false positive rate. We measured it across our early tests -- teams that checked results daily and stopped on significance had an effective false positive rate around 26%. That's not 5%. That's barely better than a coin flip between "real" and "noise."

The Bayesian replacement: what we switched to

The Bayesian approach answers a fundamentally different question. Instead of "how surprising is this data if there's no effect?" it asks "given this data, what's the probability that B is better than A?"

Here's the core of what replaced our z-test:

function bayesianTest(
  controlConversions: number, controlVisitors: number,
  variantConversions: number, variantVisitors: number,
  simulations: number = 50_000
) {
  // Beta posteriors with uniform (1,1) prior
  const alphaC = controlConversions + 1;
  const betaC = controlVisitors - controlConversions + 1;
  const alphaV = variantConversions + 1;
  const betaV = variantVisitors - variantConversions + 1;

  let variantWins = 0;
  for (let i = 0; i < simulations; i++) {
    const sampleC = betaSample(alphaC, betaC);
    const sampleV = betaSample(alphaV, betaV);
    if (sampleV > sampleC) variantWins++;
  }

  const probVariantBetter = variantWins / simulations;
  const expectedLossControl = computeExpectedLoss(alphaC, betaC, alphaV, betaV);
  const expectedLossVariant = computeExpectedLoss(alphaV, betaV, alphaC, betaC);

  return {
    probVariantBetter,
    expectedLossControl,
    expectedLossVariant,
    credibleInterval: computeCredibleInterval(alphaC, betaC, alphaV, betaV),
  };
}

The key addition is expected loss. Instead of just saying "B is probably better," we compute the expected cost (in conversion rate points) of picking the wrong variant. When expected loss drops below a threshold (we use 0.1% of baseline), we call the test.

What actually changed

After switching our stats engine and running both methods in parallel for 6 weeks on the same experiments, here's what we measured:

  • False positive rate dropped from ~26% to ~8% (mostly because Bayesian + expected loss naturally requires more evidence before calling a winner)
  • Time to decision decreased by about 20% on average -- sounds counterintuitive, but Bayesian methods handle continuous monitoring natively, so we could call clear winners earlier without guilt
  • Stakeholder confusion dropped dramatically -- "94% chance B is better" needs zero explanation. "p = 0.038" needs a 10-minute lecture.

The honest tradeoffs

I'm not going to pretend Bayesian is strictly better. Here's what I miss about frequentist:

Frequentist advantages:

  • The error guarantees are mathematically airtight (if you follow the protocol)
  • Simpler to audit and explain to statisticians
  • Lighter computationally -- a z-test is one formula, not 50K Monte Carlo samples

Bayesian advantages:

  • You can peek at results without penalty
  • Outputs are probabilities, not p-values -- humans understand them
  • Expected loss gives you a direct decision criterion
  • You can incorporate prior knowledge (we don't, but you could)

What we got wrong at first: our initial Bayesian implementation used too few Monte Carlo simulations (10K). The probability estimates were noisy, especially for close tests. Bumping to 50K solved it, but added ~15ms to each computation. We ended up precomputing results on a schedule and storing them in D1 rather than computing on every dashboard load.

-- Precomputed results stored in our D1 database
CREATE TABLE experiment_results (
  experiment_id TEXT NOT NULL,
  variant_id TEXT NOT NULL,
  prob_best REAL,
  expected_loss REAL,
  credible_interval_low REAL,
  credible_interval_high REAL,
  computed_at INTEGER NOT NULL,
  PRIMARY KEY (experiment_id, variant_id)
);

The framework matters less than the discipline

Here's the uncomfortable truth: a well-run frequentist program will outperform a sloppy Bayesian one every time. The biggest wins came not from switching frameworks, but from forcing minimum observation periods (7 days), requiring SRM checks before analysis, and auto-computing required sample sizes.

If you're building a stats engine from scratch, go Bayesian. The developer experience is better and your users will actually understand the results. If you're already using a frequentist tool and it's working, don't switch -- spend your energy on test design and sample size discipline instead.

The real enemy isn't the wrong framework. It's underpowered tests, premature stopping, and declaring winners based on noise.

Michał Pogoda-Rosikoń

Michał Pogoda-Rosikoń

Founder

Founder of SplitMonk and bards.ai. Data scientist from Wroclaw University of Technology, specializing in NLP and machine learning. Building AI-powered tools that optimize conversions on autopilot.