FAILURE MAP
← Case archive

FA-74736 / Experiment statistics / Open access

O'Brien-Fleming type alpha spending: Boundary scales with t instead of sqrt(t) · case 01

Early looks spend far too little alpha and late looks too much.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The spending function uses z / t.

VERIFIED REPAIR

Divide the critical value by sqrt(t).

Unsuccessful approach: Multiplying by sqrt(t) spends most alpha at the first look.

Case contract

Cumulative alpha spent at information fraction t is 2 - 2 Phi(z_{1-alpha/2} / sqrt(t)) (t capped at 1). Information fractions must strictly increase from 0, otherwise return None. Return the incremental alpha spent at each look rounded to 6.

Why this case matters

Spending functions let teams peek at experiments without inflating the false positive rate.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
    nd = statistics.NormalDist()
    z = nd.inv_cdf(1 - alpha / 2)
    out = []
    prev = 0.0
    last_t = 0.0
    for t in info_fractions:
        if t <= last_t:
            return None
        t = min(t, 1.0)
        spent = 2 - 2 * nd.cdf(z / t)
        out.append(round(spent - prev, 6))
        prev = spent
        last_t = t
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
  ('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
 [('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 5', [[0.2, 0.5, 0.75, 1.0], 0.05], [1.2e-05, 0.005563, 0.018051, 0.026375]),
  ('look schedule sample 6', [[0.2], 0.1], [0.000235]),
  ('look schedule sample 7', [[1.0], 0.01], [0.01])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
  ('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
  ('look schedule sample 26', [[0.25], 0.1], [0.001003])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
  ('look schedule sample 37', [[0.25, 0.4, 0.6, 0.75], 0.05], [8.9e-05, 0.001853, 0.009455, 0.012229])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
four equally spaced looks[0.0, 8.9e-05, 0.008879, 0.041032][8.9e-05, 0.005486, 0.018051, 0.026375]Failed
single final look spends all alpha[0.05][0.05]Passed
two looks at stricter alpha[0.0, 0.01][0.00027, 0.00973]Failed
repeated fraction is invalidNoneNonePassed
decreasing fractions are invalidNoneNonePassed
overrun beyond full information[0.014287, 0.035713][0.02843, 0.02157]Failed
look schedule sample 1[1.8e-05, 0.001265, 0.008717][0.000883, 0.003095, 0.006022]Failed
look schedule sample 2[0.0, 3.9e-05, 0.000964, 0.098997][0.000235, 0.009067, 0.010707, 0.079991]Failed

SHA-256 / e381d7b97f274fd8dbafe8c576022a25d723739d0df3b43fcd1cc1347e6115ef

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
    nd = statistics.NormalDist()
    z = nd.inv_cdf(1 - alpha / 2)
    out = []
    prev = 0.0
    last_t = 0.0
    for t in info_fractions:
        if t <= last_t:
            return None
        t = min(t, 1.0)
        spent = 2 - 2 * nd.cdf(z * math.sqrt(t))
        out.append(round(spent - prev, 6))
        prev = spent
        last_t = t
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
  ('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
 [('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 5', [[0.2, 0.5, 0.75, 1.0], 0.05], [1.2e-05, 0.005563, 0.018051, 0.026375]),
  ('look schedule sample 6', [[0.2], 0.1], [0.000235]),
  ('look schedule sample 7', [[1.0], 0.01], [0.01])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
  ('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
  ('look schedule sample 26', [[0.25], 0.1], [0.001003])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
  ('look schedule sample 37', [[0.25, 0.4, 0.6, 0.75], 0.05], [8.9e-05, 0.001853, 0.009455, 0.012229])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
four equally spaced looks[0.327095, -0.161319, -0.076151, -0.039625][8.9e-05, 0.005486, 0.018051, 0.026375]Failed
single final look spends all alpha[0.05][0.05]Passed
two looks at stricter alpha[0.068548, -0.058548][0.00027, 0.00973]Failed
repeated fraction is invalidNoneNonePassed
decreasing fractions are invalidNoneNonePassed
overrun beyond full information[0.079594, -0.029594][0.02843, 0.02157]Failed
look schedule sample 1[0.046018, -0.024789, -0.011229][0.000883, 0.003095, 0.006022]Failed
look schedule sample 2[0.461974, -0.163772, -0.053408, -0.144794][0.000235, 0.009067, 0.010707, 0.079991]Failed

SHA-256 / 92bcafb189b481e3532cc4c087cab5a8929a6db177e944d5739bca4b1525c7a7

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
    nd = statistics.NormalDist()
    z = nd.inv_cdf(1 - alpha / 2)
    out = []
    prev = 0.0
    last_t = 0.0
    for t in info_fractions:
        if t <= last_t:
            return None
        t = min(t, 1.0)
        spent = 2 - 2 * nd.cdf(z / math.sqrt(t))
        out.append(round(spent - prev, 6))
        prev = spent
        last_t = t
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
  ('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
 [('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 5', [[0.2, 0.5, 0.75, 1.0], 0.05], [1.2e-05, 0.005563, 0.018051, 0.026375]),
  ('look schedule sample 6', [[0.2], 0.1], [0.000235]),
  ('look schedule sample 7', [[1.0], 0.01], [0.01])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
  ('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
  ('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
  ('look schedule sample 26', [[0.25], 0.1], [0.001003])],
 [('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
  ('single final look spends all alpha', [[1.0], 0.05], [0.05]),
  ('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
  ('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
  ('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
  ('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
  ('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
  ('look schedule sample 37', [[0.25, 0.4, 0.6, 0.75], 0.05], [8.9e-05, 0.001853, 0.009455, 0.012229])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
four equally spaced looks[8.9e-05, 0.005486, 0.018051, 0.026375][8.9e-05, 0.005486, 0.018051, 0.026375]Passed
single final look spends all alpha[0.05][0.05]Passed
two looks at stricter alpha[0.00027, 0.00973][0.00027, 0.00973]Passed
repeated fraction is invalidNoneNonePassed
decreasing fractions are invalidNoneNonePassed
overrun beyond full information[0.02843, 0.02157][0.02843, 0.02157]Passed
look schedule sample 1[0.000883, 0.003095, 0.006022][0.000883, 0.003095, 0.006022]Passed
look schedule sample 2[0.000235, 0.009067, 0.010707, 0.079991][0.000235, 0.009067, 0.010707, 0.079991]Passed

SHA-256 / 00b9af7b6d9f9e7751795f8d27399282214a0ed2fb4531813da277f4a27ab3a6

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:59.589220+00:00.

Case digest / e4999861fddc39c93a615ab864a228c150776cd57bce95baab3cafcddafa4fe4