FA-74461 / Experiment statistics / Open access
Group sequential stopping rule: Every look uses the final boundary · case 01
Early looks stop at the much lower final threshold, inflating type I error.
ROOT CAUSE
The boundary is read as boundaries[-1] at each look.
VERIFIED REPAIR
Use the boundary for the current look index.
Unsuccessful approach: Using the first boundary everywhere makes late looks too strict.
Case contract
Looks beyond len(boundaries) are ignored. At look i: stop for efficacy if |z| >= boundaries[i]; else before the final planned look stop for futility if z < futility[i]; at the final planned look without efficacy return null. If data ends earlier return [continue, last look index or -1]. Return [decision, look index].
Why this case matters
Interim analyses control false positives only if the stopping rule is executed exactly.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[-1]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0]),
('interim look sample 3', [[-0.5, 2.36, 3.0, 2.36], [2.41, 2.41, 2.41], [0.5, 0.0, 0.5]], ['futility', 0])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 26', [[2.0, 1.2], [3.47, 2.45, 2.0], [0.0, 0.5, 0.5]], ['continue', 1])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 52', [[2.36, 3.0], [3.47, 2.45, 2.0], [0.5, -1.0, 0.5]], ['efficacy', 1])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 19', [[1.98, 2.01, 1.98], [3.47, 2.45, 2.0], [0.5, -1.0, 0.0]], ['null', 2])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 41', [[2.41, -0.5, 4.5, 3.0], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.0]], ['futility', 1])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['efficacy', 0] | ['continue', 0] | Failed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 3 | ['futility', 0] | ['futility', 0] | Passed |
SHA-256 / f11cf7160caaf4e59cd3cc83f35909670ea2f6e67cc53029275a19035c0feaeb
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[0]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0]),
('interim look sample 3', [[-0.5, 2.36, 3.0, 2.36], [2.41, 2.41, 2.41], [0.5, 0.0, 0.5]], ['futility', 0])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 26', [[2.0, 1.2], [3.47, 2.45, 2.0], [0.0, 0.5, 0.5]], ['continue', 1])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 52', [[2.36, 3.0], [3.47, 2.45, 2.0], [0.5, -1.0, 0.5]], ['efficacy', 1])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 19', [[1.98, 2.01, 1.98], [3.47, 2.45, 2.0], [0.5, -1.0, 0.0]], ['null', 2])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 41', [[2.41, -0.5, 4.5, 3.0], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.0]], ['futility', 1])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['continue', 1] | ['efficacy', 1] | Failed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 3 | ['futility', 0] | ['futility', 0] | Passed |
SHA-256 / 45659cc448deb1f482f25a3b8657a727bb7e4a3fb18e50d92aaa12290012ac24
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0]),
('interim look sample 3', [[-0.5, 2.36, 3.0, 2.36], [2.41, 2.41, 2.41], [0.5, 0.0, 0.5]], ['futility', 0])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 26', [[2.0, 1.2], [3.47, 2.45, 2.0], [0.0, 0.5, 0.5]], ['continue', 1])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 52', [[2.36, 3.0], [3.47, 2.45, 2.0], [0.5, -1.0, 0.5]], ['efficacy', 1])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 19', [[1.98, 2.01, 1.98], [3.47, 2.45, 2.0], [0.5, -1.0, 0.0]], ['null', 2])],
[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 41', [[2.41, -0.5, 4.5, 3.0], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.0]], ['futility', 1])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 3 | ['futility', 0] | ['futility', 0] | Passed |
SHA-256 / b1267376267dd955d71f566debd0c9d4997ffa625a699993e7faaee73f7c2268
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.098826+00:00.
Case digest / c5e17568fd1a08e9e25f46319792c3141debd914dddd7977e5677172e559b812