FAILURE MAP
← Case archive

FA-74591 / Experiment statistics / Open access

Bayesian conversion comparison: The probability sum stops one term early · case 01

Chance-to-beat-control is systematically understated.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The loop runs over range(alpha_B - 1).

VERIFIED REPAIR

Sum i from 0 to alpha_B - 1 inclusive.

Unsuccessful approach: Shifting the range to 1..alpha_B drops the first term and adds a spurious last one.

Case contract

With an integer Beta(prior_a, prior_b) prior, arm posteriors are Beta(prior_a + s, prior_b + n - s). P(B > A) uses the exact integer-parameter sum over i < alpha_B of exp(lnB(alpha_A + i, beta_A + beta_B) - ln(beta_B + i) - lnB(1 + i, beta_B) - lnB(alpha_A, beta_A)). Return [posterior mean A, posterior mean B, P(B > A)] rounded to 6.

Why this case matters

Bayesian dashboards report a "chance to beat control" that product teams act on directly.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):
    aa, ba = prior_a + a_succ, prior_b + a_n - a_succ
    ab, bb = prior_a + b_succ, prior_b + b_n - b_succ
    def lbeta(x, y):
        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)
    total = 0.0
    for i in range(ab - 1):
        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))
    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),
  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],
 [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),
  ('posterior sample 3', [8, 10, 23, 40, 1, 5], [0.5625, 0.521739, 0.384139]),
  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),
  ('posterior sample 10', [4, 5, 2, 10, 1, 1], [0.714286, 0.25, 0.017534]),
  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),
  ('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),
  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),
  ('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05]),
  ('posterior sample 24', [3, 20, 19, 20, 1, 5], [0.153846, 0.769231, 0.999999])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
uniform prior small counts[0.333333, 0.583333, 0.857942][0.333333, 0.583333, 0.90081]Failed
informative prior shifts means[0.166667, 0.25, 0.537594][0.166667, 0.25, 0.706767]Failed
equal data gives one half[0.416667, 0.416667, 0.391641][0.416667, 0.416667, 0.5]Failed
treatment clearly worse[0.727273, 0.181818, 2e-05][0.727273, 0.181818, 6.3e-05]Failed
zero successes in treatment[0.25, 0.083333, 0.0][0.25, 0.083333, 0.107143]Failed
unequal sample sizes[0.139535, 0.307692, 0.793515][0.139535, 0.307692, 0.902022]Failed
posterior sample 1[0.357143, 0.444444, 0.641358][0.357143, 0.444444, 0.724072]Failed
posterior sample 2[0.119048, 0.916667, 1.0][0.119048, 0.916667, 1.0]Passed

SHA-256 / 2ca231b0ceef7bab99eba45f6bcba9890af70a1e6a518498fa0fbbbf13bab69d

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):
    aa, ba = prior_a + a_succ, prior_b + a_n - a_succ
    ab, bb = prior_a + b_succ, prior_b + b_n - b_succ
    def lbeta(x, y):
        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)
    total = 0.0
    for i in range(1, ab + 1):
        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))
    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),
  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],
 [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),
  ('posterior sample 3', [8, 10, 23, 40, 1, 5], [0.5625, 0.521739, 0.384139]),
  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),
  ('posterior sample 10', [4, 5, 2, 10, 1, 1], [0.714286, 0.25, 0.017534]),
  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),
  ('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),
  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),
  ('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05]),
  ('posterior sample 24', [3, 20, 19, 20, 1, 5], [0.153846, 0.769231, 0.999999])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
uniform prior small counts[0.333333, 0.583333, 0.748779][0.333333, 0.583333, 0.90081]Failed
informative prior shifts means[0.166667, 0.25, 0.525172][0.166667, 0.25, 0.706767]Failed
equal data gives one half[0.416667, 0.416667, 0.539362][0.416667, 0.416667, 0.5]Failed
treatment clearly worse[0.727273, 0.181818, 0.000164][0.727273, 0.181818, 6.3e-05]Failed
zero successes in treatment[0.25, 0.083333, 0.153727][0.25, 0.083333, 0.107143]Failed
unequal sample sizes[0.139535, 0.307692, 0.66401][0.139535, 0.307692, 0.902022]Failed
posterior sample 1[0.357143, 0.444444, 0.766254][0.357143, 0.444444, 0.724072]Failed
posterior sample 2[0.119048, 0.916667, 0.119048][0.119048, 0.916667, 1.0]Failed

SHA-256 / e1d4138409670dde49a1b9dfaf0b4a86edf897ae949f2b28f5d095050f1f0a1f

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):
    aa, ba = prior_a + a_succ, prior_b + a_n - a_succ
    ab, bb = prior_a + b_succ, prior_b + b_n - b_succ
    def lbeta(x, y):
        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)
    total = 0.0
    for i in range(ab):
        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))
    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),
  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],
 [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),
  ('posterior sample 3', [8, 10, 23, 40, 1, 5], [0.5625, 0.521739, 0.384139]),
  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),
  ('posterior sample 10', [4, 5, 2, 10, 1, 1], [0.714286, 0.25, 0.017534]),
  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),
  ('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),
  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],
 [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),
  ('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05]),
  ('posterior sample 24', [3, 20, 19, 20, 1, 5], [0.153846, 0.769231, 0.999999])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
uniform prior small counts[0.333333, 0.583333, 0.90081][0.333333, 0.583333, 0.90081]Passed
informative prior shifts means[0.166667, 0.25, 0.706767][0.166667, 0.25, 0.706767]Passed
equal data gives one half[0.416667, 0.416667, 0.5][0.416667, 0.416667, 0.5]Passed
treatment clearly worse[0.727273, 0.181818, 6.3e-05][0.727273, 0.181818, 6.3e-05]Passed
zero successes in treatment[0.25, 0.083333, 0.107143][0.25, 0.083333, 0.107143]Passed
unequal sample sizes[0.139535, 0.307692, 0.902022][0.139535, 0.307692, 0.902022]Passed
posterior sample 1[0.357143, 0.444444, 0.724072][0.357143, 0.444444, 0.724072]Passed
posterior sample 2[0.119048, 0.916667, 1.0][0.119048, 0.916667, 1.0]Passed

SHA-256 / aeed50024201a73bf61dc85c76ec01894984c50f4d23ee450bc2ae821b4dd410

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.280823+00:00.

Case digest / b6e9bf0a802510a53fa25ad58e079f4b4be93e68cfcf0b369eec82f3a0172080