FA-74361 / Experiment statistics / Open access
Sample ratio mismatch check: The mismatch threshold uses the 0.05 table · case 01
Healthy experiments are routinely flagged for SRM and readouts are blocked.
ROOT CAUSE
Critical values are taken from the alpha = 0.05 row instead of alpha = 0.001.
VERIFIED REPAIR
Use the alpha = 0.001 critical values.
Unsuccessful approach: Switching to the alpha = 0.01 row is stricter but still not the specified level.
Case contract
counts and weights are per-arm lists. Expected count = total * w / sum(weights). chi2 sums (observed - expected)^2 / expected over positive-weight arms; any unit in a zero-weight arm is an immediate mismatch [None, True]. df = number of positive-weight arms - 1; mismatch iff chi2 exceeds the alpha = 0.001 critical value (10.828, 13.816, 16.266, 18.467, 20.515 for df 1..5). No units or df < 1 -> [0.0, False]. Return [round(chi2, 6), mismatch].
Why this case matters
SRM checks are the first gate on any experiment readout; a broken check hides assignment bugs.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(counts, weights):
CRIT = {1: 3.841, 2: 5.991, 3: 7.815, 4: 9.488, 5: 11.070}
total = sum(counts)
wsum = sum(weights)
if total == 0:
return [0.0, False]
chi2 = 0.0
for o, w in zip(counts, weights):
if w == 0:
if o > 0:
return [None, True]
continue
e = total * w / wsum
chi2 += (o - e) ** 2 / e
df = sum(1 for w in weights if w > 0) - 1
if df < 1:
return [0.0, False]
return [round(chi2, 6), chi2 > CRIT[df]]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),
('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),
('arm count sample 2', [[45, 53, 73, 18], [1, 50, 2, 50]], [2400.801481, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])],
[('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True]),
('arm count sample 41', [[1183, 1307, 2437], [1, 1, 2]], [6.81165, False]),
('arm count sample 59', [[0, 92, 122, 0], [3, 50, 50, 1]], [12.933832, False])],
[('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),
('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False]),
('arm count sample 33', [[68, 114, 4807], [1, 1, 50]], [11.557017, False])],
[('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False]),
('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False])],
[('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
('no units yet', [[0, 0], [1, 1]], [0.0, False]),
('arm count sample 5', [[941, 1074, 3042], [1, 1, 3]], [8.794938, False]),
('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| balanced split within noise | [0.64, False] | [0.64, False] | Passed |
| unequal design weights are respected | [0.0, False] | [0.0, False] | Passed |
| ratio weights not in percent | [0.4, False] | [0.4, False] | Passed |
| clear mismatch at alpha 0.001 | [16.0, True] | [16.0, True] | Passed |
| moderate imbalance below the strict threshold | [4.0, True] | [4.0, False] | Failed |
| arm count sample 1 | [0.263918, False] | [0.263918, False] | Passed |
| arm count sample 2 | [2400.801481, True] | [2400.801481, True] | Passed |
| arm count sample 23 | [9.423077, True] | [9.423077, False] | Failed |
SHA-256 / f2143b32ef65ecb962b864c246821f62a0381812e98b18d4ad394120f7a5bffc
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(counts, weights):
CRIT = {1: 6.635, 2: 9.210, 3: 11.345, 4: 13.277, 5: 15.086}
total = sum(counts)
wsum = sum(weights)
if total == 0:
return [0.0, False]
chi2 = 0.0
for o, w in zip(counts, weights):
if w == 0:
if o > 0:
return [None, True]
continue
e = total * w / wsum
chi2 += (o - e) ** 2 / e
df = sum(1 for w in weights if w > 0) - 1
if df < 1:
return [0.0, False]
return [round(chi2, 6), chi2 > CRIT[df]]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),
('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),
('arm count sample 2', [[45, 53, 73, 18], [1, 50, 2, 50]], [2400.801481, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])],
[('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True]),
('arm count sample 41', [[1183, 1307, 2437], [1, 1, 2]], [6.81165, False]),
('arm count sample 59', [[0, 92, 122, 0], [3, 50, 50, 1]], [12.933832, False])],
[('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),
('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False]),
('arm count sample 33', [[68, 114, 4807], [1, 1, 50]], [11.557017, False])],
[('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False]),
('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False])],
[('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
('no units yet', [[0, 0], [1, 1]], [0.0, False]),
('arm count sample 5', [[941, 1074, 3042], [1, 1, 3]], [8.794938, False]),
('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| balanced split within noise | [0.64, False] | [0.64, False] | Passed |
| unequal design weights are respected | [0.0, False] | [0.0, False] | Passed |
| ratio weights not in percent | [0.4, False] | [0.4, False] | Passed |
| clear mismatch at alpha 0.001 | [16.0, True] | [16.0, True] | Passed |
| moderate imbalance below the strict threshold | [4.0, False] | [4.0, False] | Passed |
| arm count sample 1 | [0.263918, False] | [0.263918, False] | Passed |
| arm count sample 2 | [2400.801481, True] | [2400.801481, True] | Passed |
| arm count sample 23 | [9.423077, True] | [9.423077, False] | Failed |
SHA-256 / 46e3ddad32a20343109be30b2a4cec080f513c72e23b9d1eaad27f0fb44017be
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(counts, weights):
CRIT = {1: 10.828, 2: 13.816, 3: 16.266, 4: 18.467, 5: 20.515}
total = sum(counts)
wsum = sum(weights)
if total == 0:
return [0.0, False]
chi2 = 0.0
for o, w in zip(counts, weights):
if w == 0:
if o > 0:
return [None, True]
continue
e = total * w / wsum
chi2 += (o - e) ** 2 / e
df = sum(1 for w in weights if w > 0) - 1
if df < 1:
return [0.0, False]
return [round(chi2, 6), chi2 > CRIT[df]]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),
('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),
('arm count sample 2', [[45, 53, 73, 18], [1, 50, 2, 50]], [2400.801481, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])],
[('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True]),
('arm count sample 41', [[1183, 1307, 2437], [1, 1, 2]], [6.81165, False]),
('arm count sample 59', [[0, 92, 122, 0], [3, 50, 50, 1]], [12.933832, False])],
[('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),
('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False]),
('arm count sample 33', [[68, 114, 4807], [1, 1, 50]], [11.557017, False])],
[('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False]),
('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False])],
[('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
('no units yet', [[0, 0], [1, 1]], [0.0, False]),
('arm count sample 5', [[941, 1074, 3042], [1, 1, 3]], [8.794938, False]),
('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),
('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| balanced split within noise | [0.64, False] | [0.64, False] | Passed |
| unequal design weights are respected | [0.0, False] | [0.0, False] | Passed |
| ratio weights not in percent | [0.4, False] | [0.4, False] | Passed |
| clear mismatch at alpha 0.001 | [16.0, True] | [16.0, True] | Passed |
| moderate imbalance below the strict threshold | [4.0, False] | [4.0, False] | Passed |
| arm count sample 1 | [0.263918, False] | [0.263918, False] | Passed |
| arm count sample 2 | [2400.801481, True] | [2400.801481, True] | Passed |
| arm count sample 23 | [9.423077, False] | [9.423077, False] | Passed |
SHA-256 / 1a300c3d1a8ebc1195c61446e57e8d03c0d04923306b52150893e282846e4ffd
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:55.976869+00:00.
Case digest / 85f27a462ff0f047cf7f2354a4fe2c740f78877a22b33eec60eb4297efa85bb9