FAILURE MAP
← Case archive

FA-74731 / Experiment statistics / Open access

O'Brien-Fleming type alpha spending: The spending boundary uses a one-sided quantile · case 01

The total alpha spent by the final look is double the design alpha.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

z is the 1 - alpha quantile.

VERIFIED REPAIR

Use the 1 - alpha / 2 quantile.

Unsuccessful approach: Using 1 - alpha / 4 overcorrects and underspends.

Case contract

Cumulative alpha spent at information fraction t is 2 - 2 Phi(z_{1-alpha/2} / sqrt(t)) (t capped at 1). Information fractions must strictly increase from 0, otherwise return None. Return the incremental alpha spent at each look rounded to 6.

Why this case matters

Spending functions let teams peek at experiments without inflating the false positive rate.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
    nd = statistics.NormalDist()
    z = nd.inv_cdf(1 - alpha)
    out = []
    prev = 0.0
    last_t = 0.0
    for t in info_fractions:
        if t <= last_t:
            return None
        t = min(t, 1.0)
        spent = 2 - 2 * nd.cdf(z / math.sqrt(t))
        out.append(round(spent - prev, 6))
        prev = spent
        last_t = t
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
  ('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
 [('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 4', [[0.2, 0.8], 0.05], [1.2e-05, 0.028418]),
  ('look schedule sample 6', [[0.2], 0.1], [0.000235]),
  ('look schedule sample 7', [[1.0], 0.01], [0.01])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
  ('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
  ('look schedule sample 13', [[0.25, 0.5, 0.8, 1.0], 0.1], [0.001003, 0.019006, 0.045906, 0.034085])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
  ('look schedule sample 30', [[1.0], 0.01], [0.01])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
four equally spaced looks[0.001003, 0.019006, 0.037514, 0.042477][8.9e-05, 0.005486, 0.018051, 0.026375]Failed
single final look spends all alpha[0.1][0.05]Failed
two looks at stricter alpha[0.001002, 0.018998][0.00027, 0.00973]Failed
repeated fraction is invalidNoneNonePassed
decreasing fractions are invalidNoneNonePassed
overrun beyond full information[0.065915, 0.034085][0.02843, 0.02157]Failed
look schedule sample 1[0.002671, 0.006626, 0.010703][0.000883, 0.003095, 0.006022]Failed
look schedule sample 2[0.004162, 0.038571, 0.027193, 0.130074][0.000235, 0.009067, 0.010707, 0.079991]Failed

SHA-256 / c664caa6e894e1e943139416856149c50bd004c2895b4fb5f912f9bd2620ab5a

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
    nd = statistics.NormalDist()
    z = nd.inv_cdf(1 - alpha / 4)
    out = []
    prev = 0.0
    last_t = 0.0
    for t in info_fractions:
        if t <= last_t:
            return None
        t = min(t, 1.0)
        spent = 2 - 2 * nd.cdf(z / math.sqrt(t))
        out.append(round(spent - prev, 6))
        prev = spent
        last_t = t
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
  ('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
 [('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 4', [[0.2, 0.8], 0.05], [1.2e-05, 0.028418]),
  ('look schedule sample 6', [[0.2], 0.1], [0.000235]),
  ('look schedule sample 7', [[1.0], 0.01], [0.01])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
  ('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
  ('look schedule sample 13', [[0.25, 0.5, 0.8, 1.0], 0.1], [0.001003, 0.019006, 0.045906, 0.034085])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
  ('look schedule sample 30', [[1.0], 0.01], [0.01])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
four equally spaced looks[7e-06, 0.001518, 0.008124, 0.015351][8.9e-05, 0.005486, 0.018051, 0.026375]Failed
single final look spends all alpha[0.025][0.05]Failed
two looks at stricter alpha[7.2e-05, 0.004928][0.00027, 0.00973]Failed
repeated fraction is invalidNoneNonePassed
decreasing fractions are invalidNoneNonePassed
overrun beyond full information[0.012212, 0.012788][0.02843, 0.02157]Failed
look schedule sample 1[0.00029, 0.001409, 0.003301][0.000883, 0.003095, 0.006022]Failed
look schedule sample 2[1.2e-05, 0.00193, 0.003633, 0.044425][0.000235, 0.009067, 0.010707, 0.079991]Failed

SHA-256 / d0f54dcf8bb7f486248e191c12b02da1023483f842aeae7224f53fc64a6b5262

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
    nd = statistics.NormalDist()
    z = nd.inv_cdf(1 - alpha / 2)
    out = []
    prev = 0.0
    last_t = 0.0
    for t in info_fractions:
        if t <= last_t:
            return None
        t = min(t, 1.0)
        spent = 2 - 2 * nd.cdf(z / math.sqrt(t))
        out.append(round(spent - prev, 6))
        prev = spent
        last_t = t
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
  ('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
 [('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 4', [[0.2, 0.8], 0.05], [1.2e-05, 0.028418]),
  ('look schedule sample 6', [[0.2], 0.1], [0.000235]),
  ('look schedule sample 7', [[1.0], 0.01], [0.01])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
  ('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
  ('look schedule sample 13', [[0.25, 0.5, 0.8, 1.0], 0.1], [0.001003, 0.019006, 0.045906, 0.034085])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
  ('look schedule sample 30', [[1.0], 0.01], [0.01])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
four equally spaced looks[8.9e-05, 0.005486, 0.018051, 0.026375][8.9e-05, 0.005486, 0.018051, 0.026375]Passed
single final look spends all alpha[0.05][0.05]Passed
two looks at stricter alpha[0.00027, 0.00973][0.00027, 0.00973]Passed
repeated fraction is invalidNoneNonePassed
decreasing fractions are invalidNoneNonePassed
overrun beyond full information[0.02843, 0.02157][0.02843, 0.02157]Passed
look schedule sample 1[0.000883, 0.003095, 0.006022][0.000883, 0.003095, 0.006022]Passed
look schedule sample 2[0.000235, 0.009067, 0.010707, 0.079991][0.000235, 0.009067, 0.010707, 0.079991]Passed

SHA-256 / 4c26b2efeae84b9141ab1405a4e6c4e6cb8c86d76a042386016ff5e7bd7e61ad

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:59.571215+00:00.

Case digest / 21728893c561f35dca32ac54ac5bfff86fb5685c99a13acc9c583be9fa33d377