FA-74731 / Experiment statistics / Open access
O'Brien-Fleming type alpha spending: The spending boundary uses a one-sided quantile · case 01
The total alpha spent by the final look is double the design alpha.
ROOT CAUSE
z is the 1 - alpha quantile.
VERIFIED REPAIR
Use the 1 - alpha / 2 quantile.
Unsuccessful approach: Using 1 - alpha / 4 overcorrects and underspends.
Case contract
Cumulative alpha spent at information fraction t is 2 - 2 Phi(z_{1-alpha/2} / sqrt(t)) (t capped at 1). Information fractions must strictly increase from 0, otherwise return None. Return the incremental alpha spent at each look rounded to 6.
Why this case matters
Spending functions let teams peek at experiments without inflating the false positive rate.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
nd = statistics.NormalDist()
z = nd.inv_cdf(1 - alpha)
out = []
prev = 0.0
last_t = 0.0
for t in info_fractions:
if t <= last_t:
return None
t = min(t, 1.0)
spent = 2 - 2 * nd.cdf(z / math.sqrt(t))
out.append(round(spent - prev, 6))
prev = spent
last_t = t
return out
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
[('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 4', [[0.2, 0.8], 0.05], [1.2e-05, 0.028418]),
('look schedule sample 6', [[0.2], 0.1], [0.000235]),
('look schedule sample 7', [[1.0], 0.01], [0.01])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
('look schedule sample 13', [[0.25, 0.5, 0.8, 1.0], 0.1], [0.001003, 0.019006, 0.045906, 0.034085])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
('look schedule sample 30', [[1.0], 0.01], [0.01])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| four equally spaced looks | [0.001003, 0.019006, 0.037514, 0.042477] | [8.9e-05, 0.005486, 0.018051, 0.026375] | Failed |
| single final look spends all alpha | [0.1] | [0.05] | Failed |
| two looks at stricter alpha | [0.001002, 0.018998] | [0.00027, 0.00973] | Failed |
| repeated fraction is invalid | None | None | Passed |
| decreasing fractions are invalid | None | None | Passed |
| overrun beyond full information | [0.065915, 0.034085] | [0.02843, 0.02157] | Failed |
| look schedule sample 1 | [0.002671, 0.006626, 0.010703] | [0.000883, 0.003095, 0.006022] | Failed |
| look schedule sample 2 | [0.004162, 0.038571, 0.027193, 0.130074] | [0.000235, 0.009067, 0.010707, 0.079991] | Failed |
SHA-256 / c664caa6e894e1e943139416856149c50bd004c2895b4fb5f912f9bd2620ab5a
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
nd = statistics.NormalDist()
z = nd.inv_cdf(1 - alpha / 4)
out = []
prev = 0.0
last_t = 0.0
for t in info_fractions:
if t <= last_t:
return None
t = min(t, 1.0)
spent = 2 - 2 * nd.cdf(z / math.sqrt(t))
out.append(round(spent - prev, 6))
prev = spent
last_t = t
return out
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
[('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 4', [[0.2, 0.8], 0.05], [1.2e-05, 0.028418]),
('look schedule sample 6', [[0.2], 0.1], [0.000235]),
('look schedule sample 7', [[1.0], 0.01], [0.01])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
('look schedule sample 13', [[0.25, 0.5, 0.8, 1.0], 0.1], [0.001003, 0.019006, 0.045906, 0.034085])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
('look schedule sample 30', [[1.0], 0.01], [0.01])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| four equally spaced looks | [7e-06, 0.001518, 0.008124, 0.015351] | [8.9e-05, 0.005486, 0.018051, 0.026375] | Failed |
| single final look spends all alpha | [0.025] | [0.05] | Failed |
| two looks at stricter alpha | [7.2e-05, 0.004928] | [0.00027, 0.00973] | Failed |
| repeated fraction is invalid | None | None | Passed |
| decreasing fractions are invalid | None | None | Passed |
| overrun beyond full information | [0.012212, 0.012788] | [0.02843, 0.02157] | Failed |
| look schedule sample 1 | [0.00029, 0.001409, 0.003301] | [0.000883, 0.003095, 0.006022] | Failed |
| look schedule sample 2 | [1.2e-05, 0.00193, 0.003633, 0.044425] | [0.000235, 0.009067, 0.010707, 0.079991] | Failed |
SHA-256 / d0f54dcf8bb7f486248e191c12b02da1023483f842aeae7224f53fc64a6b5262
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
import statistics
N = 1
observations = []
def solve(info_fractions, alpha):
nd = statistics.NormalDist()
z = nd.inv_cdf(1 - alpha / 2)
out = []
prev = 0.0
last_t = 0.0
for t in info_fractions:
if t <= last_t:
return None
t = min(t, 1.0)
spent = 2 - 2 * nd.cdf(z / math.sqrt(t))
out.append(round(spent - prev, 6))
prev = spent
last_t = t
return out
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 1', [[0.6, 0.8, 1.0], 0.01], [0.000883, 0.003095, 0.006022]),
('look schedule sample 2', [[0.2, 0.4, 0.5, 1.0], 0.1], [0.000235, 0.009067, 0.010707, 0.079991])],
[('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 4', [[0.2, 0.8], 0.05], [1.2e-05, 0.028418]),
('look schedule sample 6', [[0.2], 0.1], [0.000235]),
('look schedule sample 7', [[1.0], 0.01], [0.01])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 11', [[0.2, 0.25, 0.8], 0.01], [0.0, 0.0, 0.003978]),
('look schedule sample 12', [[0.5, 0.6, 0.6], 0.01], None),
('look schedule sample 13', [[0.25, 0.5, 0.8, 1.0], 0.1], [0.001003, 0.019006, 0.045906, 0.034085])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('repeated fraction is invalid', [[0.5, 0.5], 0.05], None),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 16', [[0.4, 0.5, 0.5], 0.1], None),
('look schedule sample 17', [[0.25, 0.5], 0.1], [0.001003, 0.019006]),
('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221])],
[('four equally spaced looks', [[0.25, 0.5, 0.75, 1.0], 0.05], [8.9e-05, 0.005486, 0.018051, 0.026375]),
('single final look spends all alpha', [[1.0], 0.05], [0.05]),
('two looks at stricter alpha', [[0.5, 1.0], 0.01], [0.00027, 0.00973]),
('decreasing fractions are invalid', [[0.6, 0.4], 0.05], None),
('overrun beyond full information', [[0.8, 1.2], 0.05], [0.02843, 0.02157]),
('look schedule sample 21', [[0.2, 0.4, 0.8], 0.01], [0.0, 4.6e-05, 0.003932]),
('look schedule sample 22', [[0.4, 0.75], 0.1], [0.009302, 0.048221]),
('look schedule sample 30', [[1.0], 0.01], [0.01])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| four equally spaced looks | [8.9e-05, 0.005486, 0.018051, 0.026375] | [8.9e-05, 0.005486, 0.018051, 0.026375] | Passed |
| single final look spends all alpha | [0.05] | [0.05] | Passed |
| two looks at stricter alpha | [0.00027, 0.00973] | [0.00027, 0.00973] | Passed |
| repeated fraction is invalid | None | None | Passed |
| decreasing fractions are invalid | None | None | Passed |
| overrun beyond full information | [0.02843, 0.02157] | [0.02843, 0.02157] | Passed |
| look schedule sample 1 | [0.000883, 0.003095, 0.006022] | [0.000883, 0.003095, 0.006022] | Passed |
| look schedule sample 2 | [0.000235, 0.009067, 0.010707, 0.079991] | [0.000235, 0.009067, 0.010707, 0.079991] | Passed |
SHA-256 / 4c26b2efeae84b9141ab1405a4e6c4e6cb8c86d76a042386016ff5e7bd7e61ad
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:59.571215+00:00.
Case digest / 21728893c561f35dca32ac54ac5bfff86fb5685c99a13acc9c583be9fa33d377