FA-74456 / Experiment statistics / Open access
Group sequential stopping rule: Futility can fire at the final look · case 01
Final analyses report futility instead of the null outcome.
ROOT CAUSE
The futility check is not restricted to looks before the last planned one.
VERIFIED REPAIR
Only apply futility bounds before the final planned look.
Unsuccessful approach: Making the futility comparison inclusive changes interim decisions instead.
Case contract
Looks beyond len(boundaries) are ignored. At look i: stop for efficacy if |z| >= boundaries[i]; else before the final planned look stop for futility if z < futility[i]; at the final planned look without efficacy return null. If data ends earlier return [continue, last look index or -1]. Return [decision, look index].
Why this case matters
Interim analyses control false positives only if the stopping rule is executed exactly.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1])],
[('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['futility', 1] | ['null', 1] | Failed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| futility at z exactly at bound continues | ['efficacy', 1] | ['efficacy', 1] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / 083c7159a1c735588ba30e253ee44844a3a9e8cd8608eef9030f9efedc5a43af
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z <= futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1])],
[('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| futility at z exactly at bound continues | ['futility', 0] | ['efficacy', 1] | Failed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / 1b54dbee51650edad487e955d47929143dc5a73391734a01de6837a495709db3
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1])],
[('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| futility at z exactly at bound continues | ['efficacy', 1] | ['efficacy', 1] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / f7d47a5fa797b83d1d089fe3562b050345ca7c7069bbb71312c6e73ad73886ff
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.012823+00:00.
Case digest / d48cef8cb394c33fc02e14b0ff6e9d018adc28a134154bfc5a42ee6cc2483410