FA-74471 / Experiment statistics / Open access
Group sequential stopping rule: The final look without efficacy keeps the trial running · case 01
Experiments never conclude after the last planned analysis.
ROOT CAUSE
The null decision at the final planned look is missing.
VERIFIED REPAIR
Return null at the final planned look when efficacy is not reached.
Unsuccessful approach: Declaring null only for positive final statistics leaves negative ones running.
Case contract
Looks beyond len(boundaries) are ignored. At look i: stop for efficacy if |z| >= boundaries[i]; else before the final planned look stop for futility if z < futility[i]; at the final planned look without efficacy return null. If data ends earlier return [continue, last look index or -1]. Return [decision, look index].
Why this case matters
Interim analyses control false positives only if the stopping rule is executed exactly.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 32', [[2.41, 0.0, 0.5], [2.8, 1.98], [-1.0, -1.0]], ['null', 1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
('interim look sample 13', [[1.98, 2.0, 2.36], [3.47, 2.45, 2.0], [0.5, 0.0, -1.0]], ['efficacy', 2])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
('interim look sample 16',
[[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]],
['continue', 1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['continue', 1] | ['null', 1] | Failed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| final look without efficacy is null | ['continue', 2] | ['null', 2] | Failed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / 0561e089450563cd716fe55fb9ab34b4cee1c5b596ff0ae55ae6a6f6dc54a64a
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1 and z > 0:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 32', [[2.41, 0.0, 0.5], [2.8, 1.98], [-1.0, -1.0]], ['null', 1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
('interim look sample 13', [[1.98, 2.0, 2.36], [3.47, 2.45, 2.0], [0.5, 0.0, -1.0]], ['efficacy', 2])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
('interim look sample 16',
[[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]],
['continue', 1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['continue', 1] | ['null', 1] | Failed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| final look without efficacy is null | ['null', 2] | ['null', 2] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / 43fa4e716036d1254dc08abb5880433e65964cd17345ce64ebf61576c119134b
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 32', [[2.41, 0.0, 0.5], [2.8, 1.98], [-1.0, -1.0]], ['null', 1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
('interim look sample 13', [[1.98, 2.0, 2.36], [3.47, 2.45, 2.0], [0.5, 0.0, -1.0]], ['efficacy', 2])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
('interim look sample 16',
[[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]],
['continue', 1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| final look without efficacy is null | ['null', 2] | ['null', 2] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / b8b6d40704de3f2999ae832c6a105078d6d7a2af66d04e08178be000d0b7255b
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.188938+00:00.
Case digest / 67cdce62045a6f14fa9081d1663d7491d9b2375aa43ebadfc9bff6c27b2a748f