FAILURE MAP
← Case archive

FA-74456 / Experiment statistics / Open access

Group sequential stopping rule: Futility can fire at the final look · case 01

Final analyses report futility instead of the null outcome.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The futility check is not restricted to looks before the last planned one.

VERIFIED REPAIR

Only apply futility bounds before the final planned look.

Unsuccessful approach: Making the futility comparison inclusive changes interim decisions instead.

Case contract

Looks beyond len(boundaries) are ignored. At look i: stop for efficacy if |z| >= boundaries[i]; else before the final planned look stop for futility if z < futility[i]; at the final planned look without efficacy return null. If data ends earlier return [continue, last look index or -1]. Return [decision, look index].

Why this case matters

Interim analyses control false positives only if the stopping rule is executed exactly.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(z_values, boundaries, futility):
    planned = len(boundaries)
    for i, z in enumerate(z_values[:planned]):
        if abs(z) >= boundaries[i]:
            return ['efficacy', i]
        if z < futility[i]:
            return ['futility', i]
        if i == planned - 1:
            return ['null', i]
    return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
   [[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
   ['efficacy', 1]),
  ('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 1',
   [[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
   ['efficacy', 0]),
  ('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
 [('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
  ('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
  ('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1])],
 [('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
  ('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
  ('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('final look without efficacy is null',
   [[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
   ['null', 2]),
  ('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
  ('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
efficacy exactly at the boundary['efficacy', 1]['efficacy', 1]Passed
negative effect crossing boundary stops['efficacy', 0]['efficacy', 0]Passed
futility is not applied at the final look['futility', 1]['null', 1]Failed
early look uses its own boundary['continue', 0]['continue', 0]Passed
continuing reports last look['continue', 1]['continue', 1]Passed
futility at z exactly at bound continues['efficacy', 1]['efficacy', 1]Passed
interim look sample 1['efficacy', 0]['efficacy', 0]Passed
interim look sample 2['efficacy', 0]['efficacy', 0]Passed

SHA-256 / 083c7159a1c735588ba30e253ee44844a3a9e8cd8608eef9030f9efedc5a43af

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(z_values, boundaries, futility):
    planned = len(boundaries)
    for i, z in enumerate(z_values[:planned]):
        if abs(z) >= boundaries[i]:
            return ['efficacy', i]
        if i < planned - 1 and z <= futility[i]:
            return ['futility', i]
        if i == planned - 1:
            return ['null', i]
    return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
   [[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
   ['efficacy', 1]),
  ('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 1',
   [[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
   ['efficacy', 0]),
  ('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
 [('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
  ('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
  ('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1])],
 [('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
  ('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
  ('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('final look without efficacy is null',
   [[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
   ['null', 2]),
  ('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
  ('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
efficacy exactly at the boundary['efficacy', 1]['efficacy', 1]Passed
negative effect crossing boundary stops['efficacy', 0]['efficacy', 0]Passed
futility is not applied at the final look['null', 1]['null', 1]Passed
early look uses its own boundary['continue', 0]['continue', 0]Passed
continuing reports last look['continue', 1]['continue', 1]Passed
futility at z exactly at bound continues['futility', 0]['efficacy', 1]Failed
interim look sample 1['efficacy', 0]['efficacy', 0]Passed
interim look sample 2['efficacy', 0]['efficacy', 0]Passed

SHA-256 / 1b54dbee51650edad487e955d47929143dc5a73391734a01de6837a495709db3

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(z_values, boundaries, futility):
    planned = len(boundaries)
    for i, z in enumerate(z_values[:planned]):
        if abs(z) >= boundaries[i]:
            return ['efficacy', i]
        if i < planned - 1 and z < futility[i]:
            return ['futility', i]
        if i == planned - 1:
            return ['null', i]
    return ['continue', len(z_values[:planned]) - 1]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('efficacy exactly at the boundary',
   [[1.0, 2.96], [4.33, 2.96, 2.36], [-1.0, -1.0, -1.0]],
   ['efficacy', 1]),
  ('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 1',
   [[2.41, 2.41, 2.01, 1.98], [2.41, 2.41, 2.41], [0.0, -1.0, -1.0]],
   ['efficacy', 0]),
  ('interim look sample 2', [[4.5, -2.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['efficacy', 0])],
 [('negative effect crossing boundary stops', [[-4.5], [4.33, 2.96], [0.0, 0.0]], ['efficacy', 0]),
  ('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 6', [[0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['continue', 0]),
  ('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 11', [[0.5, 0.0], [2.8, 1.98], [0.0, -1.0]], ['null', 1]),
  ('interim look sample 12', [[0.5, -0.5, -1.5], [3.47, 2.45, 2.0], [0.0, 0.5, 0.0]], ['futility', 1])],
 [('early look uses its own boundary', [[2.5], [4.33, 2.5], [-1.0, -1.0]], ['continue', 0]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('interim look sample 8', [[1.98, 2.41, -0.5], [3.47, 2.45, 2.0], [-1.0, -1.0, 0.5]], ['null', 2]),
  ('interim look sample 16', [[2.0, 2.36], [4.33, 2.96, 2.36, 2.01], [0.5, 0.5, -1.0, 0.0]], ['continue', 1]),
  ('interim look sample 17', [[3.0], [2.8, 1.98], [-1.0, 0.0]], ['efficacy', 0])],
 [('futility is not applied at the final look', [[0.5, -1.2], [2.8, 1.98], [0.0, 0.0]], ['null', 1]),
  ('continuing reports last look', [[0.5, 0.7], [4.33, 2.96, 2.36], [0.0, 0.0, 0.0]], ['continue', 1]),
  ('no looks yet', [[], [2.8, 1.98], [0.0, 0.0]], ['continue', -1]),
  ('extra looks are ignored', [[0.3, 0.2, 5.0], [2.8, 1.98], [-1.0, -1.0]], ['null', 1]),
  ('futility at z exactly at bound continues',
   [[0.0, 3.0], [2.8, 1.98, 1.5], [0.0, 0.0, 0.0]],
   ['efficacy', 1]),
  ('final look without efficacy is null',
   [[0.1, 0.2, 1.9], [3.47, 2.45, 2.0], [-1.0, -1.0, -1.0]],
   ['null', 2]),
  ('interim look sample 21', [[-1.5], [3.47, 2.45, 2.0], [-1.0, 0.5, 0.5]], ['futility', 0]),
  ('interim look sample 22', [[-2.5, -0.5], [2.8, 1.98], [0.5, -1.0]], ['futility', 0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
efficacy exactly at the boundary['efficacy', 1]['efficacy', 1]Passed
negative effect crossing boundary stops['efficacy', 0]['efficacy', 0]Passed
futility is not applied at the final look['null', 1]['null', 1]Passed
early look uses its own boundary['continue', 0]['continue', 0]Passed
continuing reports last look['continue', 1]['continue', 1]Passed
futility at z exactly at bound continues['efficacy', 1]['efficacy', 1]Passed
interim look sample 1['efficacy', 0]['efficacy', 0]Passed
interim look sample 2['efficacy', 0]['efficacy', 0]Passed

SHA-256 / f7d47a5fa797b83d1d089fe3562b050345ca7c7069bbb71312c6e73ad73886ff

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.012823+00:00.

Case digest / d48cef8cb394c33fc02e14b0ff6e9d018adc28a134154bfc5a42ee6cc2483410