FAILURE MAP
← Case archive

FA-74656 / Experiment statistics / Open access

Bootstrap difference interval: The lower tail is not halved · case 01

Lower bounds sit at twice the intended tail probability.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The lower index uses (1 - level) * reps.

VERIFIED REPAIR

Use (1 - level) / 2 * reps.

Unsuccessful approach: Rounding the halved rank up moves non-integral ranks one position inward.

Case contract

rng = random.Random(seed). Each of reps replicates resamples control then treatment with replacement (rng.choices, same sizes) and records mean(t) - mean(c). After sorting, the interval is diffs[floor((1 - level)/2 * reps)] to diffs[ceil((1 + level)/2 * reps) - 1]. Return both ends rounded to 6.

Why this case matters

Bootstrap intervals are the fallback for skewed metrics; reproducibility and indexing must be exact.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
    rng = random.Random(seed)
    diffs = []
    for _ in range(reps):
        c = rng.choices(control, k=len(control))
        t = rng.choices(treatment, k=len(treatment))
        diffs.append(sum(t) / len(t) - sum(c) / len(c))
    diffs.sort()
    lo = diffs[math.floor((1 - level) * reps)]
    hi = diffs[math.ceil((1 + level) / 2 * reps) - 1]
    return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
  ('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
  ('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
  ('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 5', [[5, 1, 3, 1, 5], [8, 0, 1, 5], 50, 479, 0.95], [-2.9, 3.45]),
  ('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
  ('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5]),
  ('bootstrap sample 8', [[3, 2, 8, 5, 4], [4, 3, 6, 8, 6], 101, 427, 0.9], [-1.2, 2.8])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
  ('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
  ('bootstrap sample 13', [[4, 8, 1, 1], [8, 8], 200, 78, 0.9], [2.0, 7.0]),
  ('bootstrap sample 18', [[3, 4, 9], [0, 8], 101, 541, 0.95], [-9.0, 4.666667])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
  ('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
  ('bootstrap sample 20', [[8, 6, 4, 2, 2], [0, 11, 7, 7], 50, 807, 0.95], [-2.65, 4.65]),
  ('bootstrap sample 28', [[1, 5, 0, 6], [9, 1, 6, 8, 2], 41, 604, 0.9], [-1.1, 5.65])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
  ('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
  ('bootstrap sample 27', [[2, 3, 6], [9, 11, 4, 11, 4], 200, 169, 0.8], [2.0, 6.533333]),
  ('bootstrap sample 45', [[6, 3, 3, 7], [1, 9, 6], 41, 871, 0.95], [-3.0, 4.25])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ninety percent interval[0.25, 4.5][0.0, 4.5]Failed
eighty percent interval[-1.75, 1.083333][-2.333333, 1.083333]Failed
treatment larger than control[4.833333, 7.0][4.833333, 7.0]Passed
ninety-five percent interval[0.35, 5.65][-0.2, 5.65]Failed
bootstrap sample 1[0.5, 4.95][-0.35, 4.95]Failed
bootstrap sample 2[0.0, 3.0][-0.4, 3.0]Failed
bootstrap sample 3[-1.85, 4.35][-2.15, 4.35]Failed
bootstrap sample 4[-1.666667, 7.0][-2.333333, 7.0]Failed

SHA-256 / 68026111416808fb48102834e824a1e2d5e001022f944ab71f786c6d772bc144

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
    rng = random.Random(seed)
    diffs = []
    for _ in range(reps):
        c = rng.choices(control, k=len(control))
        t = rng.choices(treatment, k=len(treatment))
        diffs.append(sum(t) / len(t) - sum(c) / len(c))
    diffs.sort()
    lo = diffs[math.ceil((1 - level) / 2 * reps)]
    hi = diffs[math.ceil((1 + level) / 2 * reps) - 1]
    return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
  ('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
  ('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
  ('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 5', [[5, 1, 3, 1, 5], [8, 0, 1, 5], 50, 479, 0.95], [-2.9, 3.45]),
  ('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
  ('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5]),
  ('bootstrap sample 8', [[3, 2, 8, 5, 4], [4, 3, 6, 8, 6], 101, 427, 0.9], [-1.2, 2.8])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
  ('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
  ('bootstrap sample 13', [[4, 8, 1, 1], [8, 8], 200, 78, 0.9], [2.0, 7.0]),
  ('bootstrap sample 18', [[3, 4, 9], [0, 8], 101, 541, 0.95], [-9.0, 4.666667])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
  ('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
  ('bootstrap sample 20', [[8, 6, 4, 2, 2], [0, 11, 7, 7], 50, 807, 0.95], [-2.65, 4.65]),
  ('bootstrap sample 28', [[1, 5, 0, 6], [9, 1, 6, 8, 2], 41, 604, 0.9], [-1.1, 5.65])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
  ('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
  ('bootstrap sample 27', [[2, 3, 6], [9, 11, 4, 11, 4], 200, 169, 0.8], [2.0, 6.533333]),
  ('bootstrap sample 45', [[6, 3, 3, 7], [1, 9, 6], 41, 871, 0.95], [-3.0, 4.25])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ninety percent interval[0.0, 4.5][0.0, 4.5]Passed
eighty percent interval[-2.0, 1.083333][-2.333333, 1.083333]Failed
treatment larger than control[4.833333, 7.0][4.833333, 7.0]Passed
ninety-five percent interval[0.15, 5.65][-0.2, 5.65]Failed
bootstrap sample 1[-0.25, 4.95][-0.35, 4.95]Failed
bootstrap sample 2[-0.4, 3.0][-0.4, 3.0]Passed
bootstrap sample 3[-1.9, 4.35][-2.15, 4.35]Failed
bootstrap sample 4[-2.333333, 7.0][-2.333333, 7.0]Passed

SHA-256 / be34ae4f1d00d917ded2febf2a533fffc42a8e29ce35eb51c8ea7c4b53494dce

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
    rng = random.Random(seed)
    diffs = []
    for _ in range(reps):
        c = rng.choices(control, k=len(control))
        t = rng.choices(treatment, k=len(treatment))
        diffs.append(sum(t) / len(t) - sum(c) / len(c))
    diffs.sort()
    lo = diffs[math.floor((1 - level) / 2 * reps)]
    hi = diffs[math.ceil((1 + level) / 2 * reps) - 1]
    return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
  ('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
  ('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
  ('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 5', [[5, 1, 3, 1, 5], [8, 0, 1, 5], 50, 479, 0.95], [-2.9, 3.45]),
  ('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
  ('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5]),
  ('bootstrap sample 8', [[3, 2, 8, 5, 4], [4, 3, 6, 8, 6], 101, 427, 0.9], [-1.2, 2.8])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
  ('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
  ('bootstrap sample 13', [[4, 8, 1, 1], [8, 8], 200, 78, 0.9], [2.0, 7.0]),
  ('bootstrap sample 18', [[3, 4, 9], [0, 8], 101, 541, 0.95], [-9.0, 4.666667])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
  ('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
  ('bootstrap sample 20', [[8, 6, 4, 2, 2], [0, 11, 7, 7], 50, 807, 0.95], [-2.65, 4.65]),
  ('bootstrap sample 28', [[1, 5, 0, 6], [9, 1, 6, 8, 2], 41, 604, 0.9], [-1.1, 5.65])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
  ('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
  ('bootstrap sample 27', [[2, 3, 6], [9, 11, 4, 11, 4], 200, 169, 0.8], [2.0, 6.533333]),
  ('bootstrap sample 45', [[6, 3, 3, 7], [1, 9, 6], 41, 871, 0.95], [-3.0, 4.25])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ninety percent interval[0.0, 4.5][0.0, 4.5]Passed
eighty percent interval[-2.333333, 1.083333][-2.333333, 1.083333]Passed
treatment larger than control[4.833333, 7.0][4.833333, 7.0]Passed
ninety-five percent interval[-0.2, 5.65][-0.2, 5.65]Passed
bootstrap sample 1[-0.35, 4.95][-0.35, 4.95]Passed
bootstrap sample 2[-0.4, 3.0][-0.4, 3.0]Passed
bootstrap sample 3[-2.15, 4.35][-2.15, 4.35]Passed
bootstrap sample 4[-2.333333, 7.0][-2.333333, 7.0]Passed

SHA-256 / cbd2eaaf899aae9611fb68924b1748c3e5fa2a532f3a4b363561aedb2dbdae7a

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.839473+00:00.

Case digest / 92beaac8fd0b1ace25c2f8abf35dae9bd0f36086d8602cb526908283f10ee4d2