FAILURE MAP
← Case archive

FA-74466 / Experiment statistics / Open access

Group sequential stopping rule: Continuing trials report the next look number · case 01

Monitoring shows one more completed look than actually happened.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The continue branch reports the count of looks instead of the last index.

VERIFIED REPAIR

Report the zero-based index of the last analysed look.

Unsuccessful approach: Reporting the final planned index claims all looks are done.

Case contract

Looks beyond len(boundaries) are ignored. At look i: stop for efficacy if |z| >= boundaries[i]; else before the final planned look stop for futility if z < futility[i]; at the final planned look without efficacy return null. If data ends earlier return [continue, last look index or -1]. Return [decision, look index].

Why this case matters

Interim analyses control false positives only if the stopping rule is executed exactly.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(z_values, boundaries, futility):
    planned = len(boundaries)
    for i, z in enumerate(z_values[:planned]):
        if abs(z) >= boundaries[i]:
            return ['efficacy', i]
        if i < planned - 1 and z < futility[i]:
            return ['futility', i]
        if i == planned - 1:
            return ['null', i]
    return ['continue', len(z_values)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
   [[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
   ['efficacy', 1]),
  ('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('interim look sample 1',
   [[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
   ['efficacy', 0]),
  ('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
 [('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
  ('interim look sample 7', [[], [4.33, 2.96, 2.36, 2.01], [0.0, 0.0, 0.0, -1.0]], ['continue', -1]),
  ('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
  ('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
  ('interim look sample 29', [[2.36], [2.8, 1.98], [0.5, -1.0]], ['continue', 0])],
 [('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
  ('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0]),
  ('interim look sample 18', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1])],
 [('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('final look without efficacy is null',
   [[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
   ['null', 2]),
  ('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1]),
  ('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
  ('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
efficacy exactly at the boundary['efficacy', 1]['efficacy', 1]Passed
negative effect crossing boundary stops['efficacy', 0]['efficacy', 0]Passed
futility is not applied at the final look['null', 1]['null', 1]Passed
early look uses its own boundary['continue', 1]['continue', 0]Failed
continuing reports last look['continue', 2]['continue', 1]Failed
no looks yet['continue', 0]['continue', -1]Failed
interim look sample 1['efficacy', 0]['efficacy', 0]Passed
interim look sample 2['efficacy', 0]['efficacy', 0]Passed

SHA-256 / 5baeeea1f7c4b55cc67870deb59f1002e79c14fb44c4d76588905e2bc5b95ac2

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(z_values, boundaries, futility):
    planned = len(boundaries)
    for i, z in enumerate(z_values[:planned]):
        if abs(z) >= boundaries[i]:
            return ['efficacy', i]
        if i < planned - 1 and z < futility[i]:
            return ['futility', i]
        if i == planned - 1:
            return ['null', i]
    return ['continue', planned - 1]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
   [[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
   ['efficacy', 1]),
  ('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('interim look sample 1',
   [[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
   ['efficacy', 0]),
  ('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
 [('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
  ('interim look sample 7', [[], [4.33, 2.96, 2.36, 2.01], [0.0, 0.0, 0.0, -1.0]], ['continue', -1]),
  ('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
  ('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
  ('interim look sample 29', [[2.36], [2.8, 1.98], [0.5, -1.0]], ['continue', 0])],
 [('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
  ('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0]),
  ('interim look sample 18', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1])],
 [('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('final look without efficacy is null',
   [[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
   ['null', 2]),
  ('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1]),
  ('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
  ('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
efficacy exactly at the boundary['efficacy', 1]['efficacy', 1]Passed
negative effect crossing boundary stops['efficacy', 0]['efficacy', 0]Passed
futility is not applied at the final look['null', 1]['null', 1]Passed
early look uses its own boundary['continue', 1]['continue', 0]Failed
continuing reports last look['continue', 2]['continue', 1]Failed
no looks yet['continue', 1]['continue', -1]Failed
interim look sample 1['efficacy', 0]['efficacy', 0]Passed
interim look sample 2['efficacy', 0]['efficacy', 0]Passed

SHA-256 / b6c1d5b02b79e5d18c0568f5709e078df1c015845ad3c60867048f9540a8013e

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(z_values, boundaries, futility):
    planned = len(boundaries)
    for i, z in enumerate(z_values[:planned]):
        if abs(z) >= boundaries[i]:
            return ['efficacy', i]
        if i < planned - 1 and z < futility[i]:
            return ['futility', i]
        if i == planned - 1:
            return ['null', i]
    return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
   [[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
   ['efficacy', 1]),
  ('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('interim look sample 1',
   [[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
   ['efficacy', 0]),
  ('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
 [('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
  ('interim look sample 7', [[], [4.33, 2.96, 2.36, 2.01], [0.0, 0.0, 0.0, -1.0]], ['continue', -1]),
  ('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
  ('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1]),
  ('interim look sample 29', [[2.36], [2.8, 1.98], [0.5, -1.0]], ['continue', 0])],
 [('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
  ('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0]),
  ('interim look sample 18', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1])],
 [('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('final look without efficacy is null',
   [[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
   ['null', 2]),
  ('interim look sample 10', [[], [2.8, 1.98], [0.5, -1.0]], ['continue', -1]),
  ('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
  ('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
efficacy exactly at the boundary['efficacy', 1]['efficacy', 1]Passed
negative effect crossing boundary stops['efficacy', 0]['efficacy', 0]Passed
futility is not applied at the final look['null', 1]['null', 1]Passed
early look uses its own boundary['continue', 0]['continue', 0]Passed
continuing reports last look['continue', 1]['continue', 1]Passed
no looks yet['continue', -1]['continue', -1]Passed
interim look sample 1['efficacy', 0]['efficacy', 0]Passed
interim look sample 2['efficacy', 0]['efficacy', 0]Passed

SHA-256 / 4d832e6c132df4b934b87283cd2895adcd3118cdf4495c452c1eee18ac3db6e5

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.144245+00:00.

Case digest / 98753173d31a6bf67462a0682da1a482f2c9ff724cb10c21af8dae806ba61286