FAILURE MAP
← Case archive

FA-74651 / Experiment statistics / Open access

Bootstrap difference interval: The upper percentile reads one position too far · case 01

Upper bounds are systematically too high.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The upper index omits the -1 conversion from a count to a position.

THE FAILURE

The upper index omits the -1 conversion from a count to a position.

Unsuccessful approach: Flooring before subtracting one lands one position low for non-integral ranks.

Case contract

rng = random.Random(seed). Each of reps replicates resamples control then treatment with replacement (rng.choices, same sizes) and records mean(t) - mean(c). After sorting, the interval is diffs[floor((1 - level)/2 * reps)] to diffs[ceil((1 + level)/2 * reps) - 1]. Return both ends rounded to 6.

Why this case matters

Bootstrap intervals are the fallback for skewed metrics; reproducibility and indexing must be exact.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
    rng = random.Random(seed)
    diffs = []
    for _ in range(reps):
        c = rng.choices(control, k=len(control))
        t = rng.choices(treatment, k=len(treatment))
        diffs.append(sum(t) / len(t) - sum(c) / len(c))
    diffs.sort()
    lo = diffs[math.floor((1 - level) / 2 * reps)]
    hi = diffs[math.ceil((1 + level) / 2 * reps)]
    return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
  ('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
  ('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
  ('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
  ('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5]),
  ('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
  ('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
  ('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
  ('bootstrap sample 24', [[1, 1, 3], [1, 8], 50, 923, 0.95], [-1.333333, 6.333333]),
  ('bootstrap sample 32', [[2, 6, 6, 7], [3, 1, 2], 50, 938, 0.9], [-4.75, -1.666667])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
  ('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
  ('bootstrap sample 35', [[2, 0, 5, 8, 2], [5, 10, 2, 10], 50, 323, 0.8], [0.95, 5.75]),
  ('bootstrap sample 52', [[1, 4, 4, 3, 1], [5, 9, 8], 101, 498, 0.95], [2.6, 6.733333])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 9', [[2, 8, 6], [5, 2, 1, 11, 0], 50, 416, 0.9], [-5.266667, 3.6]),
  ('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
  ('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
  ('bootstrap sample 50', [[4, 4, 9, 1, 8], [2, 1, 2, 0], 101, 193, 0.8], [-5.75, -2.1])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ninety percent interval[0.0, 4.5][0.0, 4.5]Passed
eighty percent interval[-2.333333, 1.083333][-2.333333, 1.083333]Passed
treatment larger than control[4.833333, 7.333333][4.833333, 7.0]Failed
ninety-five percent interval[-0.2, 5.9][-0.2, 5.65]Failed
bootstrap sample 1[-0.35, 4.95][-0.35, 4.95]Passed
bootstrap sample 2[-0.4, 3.0][-0.4, 3.0]Passed
bootstrap sample 3[-2.15, 4.7][-2.15, 4.35]Failed
bootstrap sample 4[-2.333333, 7.0][-2.333333, 7.0]Passed

SHA-256 / cb79fae55022eb9d16d35ff302c2af4c6f810f0dd4777d884bd496e032308178

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import random
N = 1
observations = []
def solve(control, treatment, reps, seed, level):
    rng = random.Random(seed)
    diffs = []
    for _ in range(reps):
        c = rng.choices(control, k=len(control))
        t = rng.choices(treatment, k=len(treatment))
        diffs.append(sum(t) / len(t) - sum(c) / len(c))
    diffs.sort()
    lo = diffs[math.floor((1 - level) / 2 * reps)]
    hi = diffs[math.floor((1 + level) / 2 * reps) - 1]
    return [round(lo, 6), round(hi, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 1', [[6, 3, 2, 7], [7, 9, 2, 9, 7], 200, 662, 0.9], [-0.35, 4.95]),
  ('bootstrap sample 2', [[4, 4, 4, 8, 2], [10, 9, 6, 3, 2], 50, 548, 0.8], [-0.4, 3.0]),
  ('bootstrap sample 3', [[6, 9, 2, 1], [3, 3, 11, 8, 2], 50, 269, 0.9], [-2.15, 4.35]),
  ('bootstrap sample 4', [[4, 6, 4], [3, 11], 200, 10, 0.9], [-2.333333, 7.0])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 6', [[9, 8], [11, 10, 6], 50, 172, 0.95], [-3.0, 2.5]),
  ('bootstrap sample 7', [[5, 7, 2, 9], [1, 10], 41, 416, 0.9], [-6.0, 5.5]),
  ('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
  ('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 11', [[7, 5, 6, 7, 6], [6, 4, 6, 3], 41, 117, 0.9], [-2.9, -0.3]),
  ('bootstrap sample 12', [[3, 5, 5, 0, 4], [5, 8], 41, 271, 0.95], [1.0, 5.6]),
  ('bootstrap sample 24', [[1, 1, 3], [1, 8], 50, 923, 0.95], [-1.333333, 6.333333]),
  ('bootstrap sample 32', [[2, 6, 6, 7], [3, 1, 2], 50, 938, 0.9], [-4.75, -1.666667])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 16', [[3, 4], [5, 7, 7, 7], 101, 604, 0.95], [2.0, 4.0]),
  ('bootstrap sample 17', [[0, 1], [10, 2, 5, 8, 11], 50, 644, 0.8], [4.5, 8.7]),
  ('bootstrap sample 35', [[2, 0, 5, 8, 2], [5, 10, 2, 10], 50, 323, 0.8], [0.95, 5.75]),
  ('bootstrap sample 52', [[1, 4, 4, 3, 1], [5, 9, 8], 101, 498, 0.95], [2.6, 6.733333])],
 [('ninety percent interval', [[1, 2, 3, 4], [2, 3, 5, 8], 101, 7, 0.9], [0.0, 4.5]),
  ('eighty percent interval', [[0, 0, 1, 5], [1, 1, 2], 50, 11, 0.8], [-2.333333, 1.083333]),
  ('treatment larger than control', [[1, 2], [6, 9, 7], 41, 3, 0.9], [4.833333, 7.0]),
  ('ninety-five percent interval', [[3, 1, 4, 1, 5], [9, 2, 6, 5], 200, 42, 0.95], [-0.2, 5.65]),
  ('bootstrap sample 9', [[2, 8, 6], [5, 2, 1, 11, 0], 50, 416, 0.9], [-5.266667, 3.6]),
  ('bootstrap sample 21', [[2, 6, 8, 0, 8], [9, 10, 5, 7, 10], 41, 153, 0.8], [1.2, 5.2]),
  ('bootstrap sample 22', [[8, 3], [0, 7, 11, 7, 1], 99, 38, 0.8], [-4.0, 3.4]),
  ('bootstrap sample 50', [[4, 4, 9, 1, 8], [2, 1, 2, 0], 101, 193, 0.8], [-5.75, -2.1])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ninety percent interval[0.0, 4.25][0.0, 4.5]Failed
eighty percent interval[-2.333333, 1.083333][-2.333333, 1.083333]Passed
treatment larger than control[4.833333, 7.0][4.833333, 7.0]Passed
ninety-five percent interval[-0.2, 5.65][-0.2, 5.65]Passed
bootstrap sample 1[-0.35, 4.95][-0.35, 4.95]Passed
bootstrap sample 2[-0.4, 3.0][-0.4, 3.0]Passed
bootstrap sample 3[-2.15, 3.95][-2.15, 4.35]Failed
bootstrap sample 4[-2.333333, 7.0][-2.333333, 7.0]Passed

SHA-256 / 9a4ddd2252fb5aa56dc157879a8c1b95f3904da97f5905acc56b2be1fca498d8

HELD IN THE MEMBER ARCHIVE

The verified repair and its recorded checks are member-only.

This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.

Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.

Member access is invitation-based. Sign in with your invited account to inspect the repair.

Sign in to the archive ↗

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.782686+00:00.

Case digest / f52ff2a15657a9ca344e169e748d5038342bcacc9eb13f3c2a3a895afd0f63bb