FA-74466 / Experiment statistics / Open access
Group sequential stopping rule: Continuing trials report the next look number · case 01
Monitoring shows one more completed look than actually happened.
ROOT CAUSE
The continue branch reports the count of looks instead of the last index.
VERIFIED REPAIR
Report the zero-based index of the last analysed look.
Unsuccessful approach: Reporting the final planned index claims all looks are done.
Case contract
Looks beyond len(boundaries) are ignored. At look i: stop for efficacy if |z| >= boundaries[i]; else before the final planned look stop for futility if z < futility[i]; at the final planned look without efficacy return null. If data ends earlier return [continue, last look index or -1]. Return [decision, look index].
Why this case matters
Interim analyses control false positives only if the stopping rule is executed exactly.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 7', [[], [4.33, 2.96, 2.36, 2.01], [0.0, 0.0, 0.0, -1.0]], ['continue', -1]),
('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
('interim look sample 29', [[2.36], [2.8, 1.98], [0.5, -1.0]], ['continue', 0])],
[('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0]),
('interim look sample 18', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1])],
[('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 1] | ['continue', 0] | Failed |
| continuing reports last look | ['continue', 2] | ['continue', 1] | Failed |
| no looks yet | ['continue', 0] | ['continue', -1] | Failed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / 5baeeea1f7c4b55cc67870deb59f1002e79c14fb44c4d76588905e2bc5b95ac2
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', planned - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 7', [[], [4.33, 2.96, 2.36, 2.01], [0.0, 0.0, 0.0, -1.0]], ['continue', -1]),
('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
('interim look sample 29', [[2.36], [2.8, 1.98], [0.5, -1.0]], ['continue', 0])],
[('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0]),
('interim look sample 18', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1])],
[('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 1] | ['continue', 0] | Failed |
| continuing reports last look | ['continue', 2] | ['continue', 1] | Failed |
| no looks yet | ['continue', 1] | ['continue', -1] | Failed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / b6c1d5b02b79e5d18c0568f5709e078df1c015845ad3c60867048f9540a8013e
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(z_values, boundaries, futility):
planned = len(boundaries)
for i, z in enumerate(z_values[:planned]):
if abs(z) >= boundaries[i]:
return ['efficacy', i]
if i < planned - 1 and z < futility[i]:
return ['futility', i]
if i == planned - 1:
return ['null', i]
return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
[[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
['efficacy', 1]),
('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 1',
[[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
['efficacy', 0]),
('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
[('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
('interim look sample 7', [[], [4.33, 2.96, 2.36, 2.01], [0.0, 0.0, 0.0, -1.0]], ['continue', -1]),
('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1])],
[('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
('interim look sample 29', [[2.36], [2.8, 1.98], [0.5, -1.0]], ['continue', 0])],
[('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0]),
('interim look sample 18', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1])],
[('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
('futility at z exactly at bound continues',
[[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
['efficacy', 1]),
('final look without efficacy is null',
[[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
['null', 2]),
('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1]),
('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| efficacy exactly at the boundary | ['efficacy', 1] | ['efficacy', 1] | Passed |
| negative effect crossing boundary stops | ['efficacy', 0] | ['efficacy', 0] | Passed |
| futility is not applied at the final look | ['null', 1] | ['null', 1] | Passed |
| early look uses its own boundary | ['continue', 0] | ['continue', 0] | Passed |
| continuing reports last look | ['continue', 1] | ['continue', 1] | Passed |
| no looks yet | ['continue', -1] | ['continue', -1] | Passed |
| interim look sample 1 | ['efficacy', 0] | ['efficacy', 0] | Passed |
| interim look sample 2 | ['efficacy', 0] | ['efficacy', 0] | Passed |
SHA-256 / 4d832e6c132df4b934b87283cd2895adcd3118cdf4495c452c1eee18ac3db6e5
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.144245+00:00.
Case digest / 98753173d31a6bf67462a0682da1a482f2c9ff724cb10c21af8dae806ba61286