FA-74666 / Experiment statistics / Open access
Bootstrap difference interval: Treatment is resampled before control · case 01
Intervals are not reproducible against the reference implementation for the same seed.
ROOT CAUSE
The two resampling calls are swapped, consuming the random stream in a different order.
VERIFIED REPAIR
Draw the control resample first, then the treatment resample.
Unsuccessful approach: Reseeding inside the loop makes every replicate identical.
Case contract
rng = random.Random(seed). Each of reps replicates resamples control then treatment with replacement (rng.choices, same sizes) and records mean(t) - mean(c). After sorting, the interval is diffs[floor((1 - level)/2 * reps)] to diffs[ceil((1 + level)/2 * reps) - 1]. Return both ends rounded to 6.
Why this case matters
Bootstrap intervals are the fallback for skewed metrics; reproducibility and indexing must be exact.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
rng = random.Random(seed)
diffs = []
for _ in range(reps):
t = rng.choices(treatment, k=len(treatment))
c = rng.choices(control, k=len(control))
diffs.append(sum(t) / len(t) - sum(c) / len(c))
diffs.sort()
lo = diffs[math.floor((1 - level) / 2 * reps)]
hi = diffs[math.ceil((1 + level) / 2 * reps) - 1]
return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0]),
('bootstrap sample 5', [[5, 1, 3, 1, 5], [8, 0, 1, 5], 50, 479, 0.95], [-2.9, 3.45]),
('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
('bootstrap sample 13', [[4, 8, 1, 1], [8, 8], 200, 78, 0.9], [2.0, 7.0]),
('bootstrap sample 14', [[2, 0, 8, 3], [4, 6, 7], 99, 617, 0.95], [-0.333333, 5.416667])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
('bootstrap sample 18', [[3, 4, 9], [0, 8], 101, 541, 0.95], [-9.0, 4.666667]),
('bootstrap sample 19', [[3, 3, 9], [2, 2], 99, 846, 0.9], [-7.0, -1.0])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
('bootstrap sample 25', [[5, 9, 1, 0], [7, 6, 11, 8], 41, 751, 0.8], [2.0, 6.25]),
('bootstrap sample 26', [[5, 9, 9], [7, 0, 8], 200, 180, 0.8], [-5.333333, 0.333333])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ninety percent interval | [0.0, 4.0] | [0.0, 4.5] | Failed |
| eighty percent interval | [-1.75, 1.0] | [-2.333333, 1.083333] | Failed |
| treatment larger than control | [4.666667, 7.333333] | [4.833333, 7.0] | Failed |
| ninety-five percent interval | [0.15, 5.8] | [-0.2, 5.65] | Failed |
| bootstrap sample 1 | [-0.45, 4.7] | [-0.35, 4.95] | Failed |
| bootstrap sample 2 | [-1.0, 3.2] | [-0.4, 3.0] | Failed |
| bootstrap sample 3 | [-3.3, 3.95] | [-2.15, 4.35] | Failed |
| bootstrap sample 4 | [-2.333333, 7.0] | [-2.333333, 7.0] | Passed |
SHA-256 / 7701372b15d80a47d2467c10f10141543d3939ef5fe8e12368768817bf0ae8f3
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
rng = random.Random(seed)
diffs = []
for _ in range(reps):
rng = random.Random(seed)
c = rng.choices(control, k=len(control))
t = rng.choices(treatment, k=len(treatment))
diffs.append(sum(t) / len(t) - sum(c) / len(c))
diffs.sort()
lo = diffs[math.floor((1 - level) / 2 * reps)]
hi = diffs[math.ceil((1 + level) / 2 * reps) - 1]
return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0]),
('bootstrap sample 5', [[5, 1, 3, 1, 5], [8, 0, 1, 5], 50, 479, 0.95], [-2.9, 3.45]),
('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
('bootstrap sample 13', [[4, 8, 1, 1], [8, 8], 200, 78, 0.9], [2.0, 7.0]),
('bootstrap sample 14', [[2, 0, 8, 3], [4, 6, 7], 99, 617, 0.95], [-0.333333, 5.416667])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
('bootstrap sample 18', [[3, 4, 9], [0, 8], 101, 541, 0.95], [-9.0, 4.666667]),
('bootstrap sample 19', [[3, 3, 9], [2, 2], 99, 846, 0.9], [-7.0, -1.0])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
('bootstrap sample 25', [[5, 9, 1, 0], [7, 6, 11, 8], 41, 751, 0.8], [2.0, 6.25]),
('bootstrap sample 26', [[5, 9, 9], [7, 0, 8], 200, 180, 0.8], [-5.333333, 0.333333])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ninety percent interval | [2.0, 2.0] | [0.0, 4.5] | Failed |
| eighty percent interval | [-0.5, -0.5] | [-2.333333, 1.083333] | Failed |
| treatment larger than control | [7.5, 7.5] | [4.833333, 7.0] | Failed |
| ninety-five percent interval | [4.1, 4.1] | [-0.2, 5.65] | Failed |
| bootstrap sample 1 | [3.3, 3.3] | [-0.35, 4.95] | Failed |
| bootstrap sample 2 | [-0.8, -0.8] | [-0.4, 3.0] | Failed |
| bootstrap sample 3 | [1.9, 1.9] | [-2.15, 4.35] | Failed |
| bootstrap sample 4 | [1.0, 1.0] | [-2.333333, 7.0] | Failed |
SHA-256 / 8f06cc384c5a982b52055c9258bea5bff171e22e3af0f78f0eb4c2a227e509e1
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
rng = random.Random(seed)
diffs = []
for _ in range(reps):
c = rng.choices(control, k=len(control))
t = rng.choices(treatment, k=len(treatment))
diffs.append(sum(t) / len(t) - sum(c) / len(c))
diffs.sort()
lo = diffs[math.floor((1 - level) / 2 * reps)]
hi = diffs[math.ceil((1 + level) / 2 * reps) - 1]
return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0]),
('bootstrap sample 5', [[5, 1, 3, 1, 5], [8, 0, 1, 5], 50, 479, 0.95], [-2.9, 3.45]),
('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
('bootstrap sample 13', [[4, 8, 1, 1], [8, 8], 200, 78, 0.9], [2.0, 7.0]),
('bootstrap sample 14', [[2, 0, 8, 3], [4, 6, 7], 99, 617, 0.95], [-0.333333, 5.416667])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
('bootstrap sample 18', [[3, 4, 9], [0, 8], 101, 541, 0.95], [-9.0, 4.666667]),
('bootstrap sample 19', [[3, 3, 9], [2, 2], 99, 846, 0.9], [-7.0, -1.0])],
[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
('bootstrap sample 25', [[5, 9, 1, 0], [7, 6, 11, 8], 41, 751, 0.8], [2.0, 6.25]),
('bootstrap sample 26', [[5, 9, 9], [7, 0, 8], 200, 180, 0.8], [-5.333333, 0.333333])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ninety percent interval | [0.0, 4.5] | [0.0, 4.5] | Passed |
| eighty percent interval | [-2.333333, 1.083333] | [-2.333333, 1.083333] | Passed |
| treatment larger than control | [4.833333, 7.0] | [4.833333, 7.0] | Passed |
| ninety-five percent interval | [-0.2, 5.65] | [-0.2, 5.65] | Passed |
| bootstrap sample 1 | [-0.35, 4.95] | [-0.35, 4.95] | Passed |
| bootstrap sample 2 | [-0.4, 3.0] | [-0.4, 3.0] | Passed |
| bootstrap sample 3 | [-2.15, 4.35] | [-2.15, 4.35] | Passed |
| bootstrap sample 4 | [-2.333333, 7.0] | [-2.333333, 7.0] | Passed |
SHA-256 / b8c8be71f84f96939d099b1d224f6cf4eb7720e0f512192cecfae5b7a1788766
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.926541+00:00.
Case digest / ca2d0ac5aea2c9ebb390ee89d5d35b266dbce8290ecf5ea8bc16d89232f7396a