FA-74586 / Experiment statistics / Open access
Bayesian conversion comparison: Posterior means report raw conversion rates · case 01
Displayed means ignore the prior, contradicting the reported probability.
ROOT CAUSE
The means are s / n rather than alpha / (alpha + beta).
VERIFIED REPAIR
Report alpha / (alpha + beta) of each posterior.
Unsuccessful approach: The posterior mode (alpha - 1) / (alpha + beta - 2) is a different summary.
Case contract
With an integer Beta(prior_a, prior_b) prior, arm posteriors are Beta(prior_a + s, prior_b + n - s). P(B > A) uses the exact integer-parameter sum over i < alpha_B of exp(lnB(alpha_A + i, beta_A + beta_B) - ln(beta_B + i) - lnB(1 + i, beta_B) - lnB(alpha_A, beta_A)). Return [posterior mean A, posterior mean B, P(B > A)] rounded to 6.
Why this case matters
Bayesian dashboards report a "chance to beat control" that product teams act on directly.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):
aa, ba = prior_a + a_succ, prior_b + a_n - a_succ
ab, bb = prior_a + b_succ, prior_b + b_n - b_succ
def lbeta(x, y):
return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)
total = 0.0
for i in range(ab):
total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))
return [round(a_succ / a_n, 6), round(b_succ / b_n, 6), round(total, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),
('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],
[('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),
('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),
('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),
('posterior sample 10', [4, 5, 2, 10, 1, 1], [0.714286, 0.25, 0.017534]),
('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),
('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),
('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),
('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05]),
('posterior sample 24', [3, 20, 19, 20, 1, 5], [0.153846, 0.769231, 0.999999])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| uniform prior small counts | [0.3, 0.6, 0.90081] | [0.333333, 0.583333, 0.90081] | Failed |
| informative prior shifts means | [0.0, 0.2, 0.706767] | [0.166667, 0.25, 0.706767] | Failed |
| equal data gives one half | [0.4, 0.4, 0.5] | [0.416667, 0.416667, 0.5] | Failed |
| treatment clearly worse | [0.75, 0.15, 6.3e-05] | [0.727273, 0.181818, 6.3e-05] | Failed |
| zero successes in treatment | [0.2, 0.0, 0.107143] | [0.25, 0.083333, 0.107143] | Failed |
| unequal sample sizes | [0.125, 0.3, 0.902022] | [0.139535, 0.307692, 0.902022] | Failed |
| posterior sample 1 | [0.35, 0.5, 0.724072] | [0.357143, 0.444444, 0.724072] | Failed |
| posterior sample 2 | [0.1, 1.0, 1.0] | [0.119048, 0.916667, 1.0] | Failed |
SHA-256 / 0cf60e8603202777b8ceef45da1a77c717e5f0c6a6a4044102b0ff1e05d50461
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):
aa, ba = prior_a + a_succ, prior_b + a_n - a_succ
ab, bb = prior_a + b_succ, prior_b + b_n - b_succ
def lbeta(x, y):
return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)
total = 0.0
for i in range(ab):
total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))
return [round((aa - 1) / (aa + ba - 2), 6), round((ab - 1) / (ab + bb - 2), 6), round(total, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),
('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],
[('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),
('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),
('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),
('posterior sample 10', [4, 5, 2, 10, 1, 1], [0.714286, 0.25, 0.017534]),
('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),
('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),
('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),
('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05]),
('posterior sample 24', [3, 20, 19, 20, 1, 5], [0.153846, 0.769231, 0.999999])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| uniform prior small counts | [0.3, 0.6, 0.90081] | [0.333333, 0.583333, 0.90081] | Failed |
| informative prior shifts means | [0.1, 0.2, 0.706767] | [0.166667, 0.25, 0.706767] | Failed |
| equal data gives one half | [0.4, 0.4, 0.5] | [0.416667, 0.416667, 0.5] | Failed |
| treatment clearly worse | [0.75, 0.15, 6.3e-05] | [0.727273, 0.181818, 6.3e-05] | Failed |
| zero successes in treatment | [0.2, 0.0, 0.107143] | [0.25, 0.083333, 0.107143] | Failed |
| unequal sample sizes | [0.121951, 0.272727, 0.902022] | [0.139535, 0.307692, 0.902022] | Failed |
| posterior sample 1 | [0.346154, 0.4375, 0.724072] | [0.357143, 0.444444, 0.724072] | Failed |
| posterior sample 2 | [0.1, 1.0, 1.0] | [0.119048, 0.916667, 1.0] | Failed |
SHA-256 / 0c7d8122c6dd4bdbc0eca06f3c1c2dedd45a71ea5690c100184ba75d5c990184
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):
aa, ba = prior_a + a_succ, prior_b + a_n - a_succ
ab, bb = prior_a + b_succ, prior_b + b_n - b_succ
def lbeta(x, y):
return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)
total = 0.0
for i in range(ab):
total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))
return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),
('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],
[('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),
('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),
('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),
('posterior sample 10', [4, 5, 2, 10, 1, 1], [0.714286, 0.25, 0.017534]),
('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),
('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),
('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],
[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),
('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),
('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),
('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),
('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),
('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),
('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05]),
('posterior sample 24', [3, 20, 19, 20, 1, 5], [0.153846, 0.769231, 0.999999])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| uniform prior small counts | [0.333333, 0.583333, 0.90081] | [0.333333, 0.583333, 0.90081] | Passed |
| informative prior shifts means | [0.166667, 0.25, 0.706767] | [0.166667, 0.25, 0.706767] | Passed |
| equal data gives one half | [0.416667, 0.416667, 0.5] | [0.416667, 0.416667, 0.5] | Passed |
| treatment clearly worse | [0.727273, 0.181818, 6.3e-05] | [0.727273, 0.181818, 6.3e-05] | Passed |
| zero successes in treatment | [0.25, 0.083333, 0.107143] | [0.25, 0.083333, 0.107143] | Passed |
| unequal sample sizes | [0.139535, 0.307692, 0.902022] | [0.139535, 0.307692, 0.902022] | Passed |
| posterior sample 1 | [0.357143, 0.444444, 0.724072] | [0.357143, 0.444444, 0.724072] | Passed |
| posterior sample 2 | [0.119048, 0.916667, 1.0] | [0.119048, 0.916667, 1.0] | Passed |
SHA-256 / 7562d03062cad527c877e34d18a3e0802d01f832eab1b6a98099c38bdfea8e2e
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.198334+00:00.
Case digest / 81fcaf44dcc996f8c6f89c8ef26259df234cf3603217131bfc6cb677f00e76ea