FAILURE MAP
← Case archive

FA-74491 / Experiment statistics / Open access

Holm correction across metrics: Results are reported in sorted order · case 01

Rejection flags attach to the wrong metric names.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

Results are stored at the sorted rank rather than the original index.

THE FAILURE

Results are stored at the sorted rank rather than the original index.

Unsuccessful approach: Mapping only rejections back leaves adjusted p-values misaligned.

Case contract

Holm step-down: sort p-values ascending (ties by position); reject while p_(k) <= alpha / (m - k) for k = 0, 1, ... and stop at the first failure. Adjusted p_(k) = running max of min(1, (m - k) p_(k)). Return [rejections, adjusted p-values rounded to 6] in the original metric order.

Why this case matters

Experiments track many metrics; multiplicity control keeps secondary wins honest.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(pvalues, alpha):
    m = len(pvalues)
    order = sorted(range(m), key=lambda i: (pvalues[i], i))
    reject = [False] * m
    adjusted = [0.0] * m
    running = 0.0
    stopped = False
    for rank, i in enumerate(order):
        running = max(running, min(1.0, (m - rank) * pvalues[i]))
        adjusted[rank] = round(running, 6)
        if not stopped and pvalues[i] <= alpha / (m - rank):
            reject[rank] = True
        else:
            stopped = True
    return [reject, adjusted]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 1',
   [[0.025, 0.02, 0.05, 0.6], 0.05],
   [[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]),
  ('metric p-value sample 2',
   [[0.0125, 0.03, 0.0167], 0.05],
   [[True, True, True], [0.0375, 0.0375, 0.0375]])],
 [('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 6', [[0.01, 0.05], 0.05], [[True, True], [0.02, 0.05]]),
  ('metric p-value sample 7', [[0.02, 0.0167], 0.05], [[True, True], [0.0334, 0.0334]]),
  ('metric p-value sample 17', [[0.2, 0.02, 0.05], 0.05], [[False, False, False], [0.2, 0.06, 0.1]])],
 [('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 12',
   [[0.04, 0.025, 0.01, 0.0167, 0.0125], 0.05],
   [[False, False, True, False, True], [0.0501, 0.0501, 0.05, 0.0501, 0.05]]),
  ('metric p-value sample 29',
   [[0.02, 0.2, 0.03, 0.01], 0.05],
   [[False, False, False, True], [0.06, 0.2, 0.06, 0.04]])],
 [('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 16',
   [[0.03, 0.02, 0.2, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.09, 0.0835, 0.4, 0.6, 0.0835]]),
  ('metric p-value sample 17', [[0.2, 0.02, 0.05], 0.05], [[False, False, False], [0.2, 0.06, 0.1]]),
  ('metric p-value sample 45', [[0.6, 0.0125], 0.05], [[False, True], [0.6, 0.025]])],
 [('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 21',
   [[0.0167, 0.6, 0.2, 0.02, 0.01], 0.05],
   [[False, False, False, False, True], [0.0668, 0.6, 0.4, 0.0668, 0.05]]),
  ('metric p-value sample 22', [[0.0167], 0.05], [[True], [0.0167]]),
  ('metric p-value sample 23',
   [[0.2, 0.2, 0.02, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.6, 0.6, 0.0835, 0.6, 0.0835]])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
all three pass their step thresholds[[True, True, True], [0.03, 0.03, 0.04]][[True, True, True], [0.03, 0.04, 0.03]]Failed
step-down stops at first non-rejection[[False, False, False], [0.06, 0.06, 0.06]][[False, False, False], [0.06, 0.06, 0.06]]Passed
boundary p equals alpha over remaining[[True, False, False, False], [0.05, 1.0, 1.0, 1.0]][[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]Passed
adjusted values are monotone[[True, True, False], [0.03, 0.03, 0.5]][[True, True, False], [0.03, 0.03, 0.5]]Passed
adjusted values cap at one[[False, False], [0.8, 0.8]][[False, False], [0.8, 0.8]]Passed
original order is preserved[[True, True, False], [0.003, 0.04, 0.3]][[False, True, True], [0.3, 0.003, 0.04]]Failed
metric p-value sample 1[[False, False, False, False], [0.08, 0.08, 0.1, 0.6]][[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]Passed
metric p-value sample 2[[True, True, True], [0.0375, 0.0375, 0.0375]][[True, True, True], [0.0375, 0.0375, 0.0375]]Passed

SHA-256 / 5b018692ee596eb4bf91ed0d30e2049c3840fbf2a317d6198f2e8fa5e153519c

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(pvalues, alpha):
    m = len(pvalues)
    order = sorted(range(m), key=lambda i: (pvalues[i], i))
    reject = [False] * m
    adjusted = [0.0] * m
    running = 0.0
    stopped = False
    for rank, i in enumerate(order):
        running = max(running, min(1.0, (m - rank) * pvalues[i]))
        adjusted[rank] = round(running, 6)
        if not stopped and pvalues[i] <= alpha / (m - rank):
            reject[i] = True
        else:
            stopped = True
    return [reject, adjusted]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 1',
   [[0.025, 0.02, 0.05, 0.6], 0.05],
   [[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]),
  ('metric p-value sample 2',
   [[0.0125, 0.03, 0.0167], 0.05],
   [[True, True, True], [0.0375, 0.0375, 0.0375]])],
 [('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 6', [[0.01, 0.05], 0.05], [[True, True], [0.02, 0.05]]),
  ('metric p-value sample 7', [[0.02, 0.0167], 0.05], [[True, True], [0.0334, 0.0334]]),
  ('metric p-value sample 17', [[0.2, 0.02, 0.05], 0.05], [[False, False, False], [0.2, 0.06, 0.1]])],
 [('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 12',
   [[0.04, 0.025, 0.01, 0.0167, 0.0125], 0.05],
   [[False, False, True, False, True], [0.0501, 0.0501, 0.05, 0.0501, 0.05]]),
  ('metric p-value sample 29',
   [[0.02, 0.2, 0.03, 0.01], 0.05],
   [[False, False, False, True], [0.06, 0.2, 0.06, 0.04]])],
 [('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 16',
   [[0.03, 0.02, 0.2, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.09, 0.0835, 0.4, 0.6, 0.0835]]),
  ('metric p-value sample 17', [[0.2, 0.02, 0.05], 0.05], [[False, False, False], [0.2, 0.06, 0.1]]),
  ('metric p-value sample 45', [[0.6, 0.0125], 0.05], [[False, True], [0.6, 0.025]])],
 [('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 21',
   [[0.0167, 0.6, 0.2, 0.02, 0.01], 0.05],
   [[False, False, False, False, True], [0.0668, 0.6, 0.4, 0.0668, 0.05]]),
  ('metric p-value sample 22', [[0.0167], 0.05], [[True], [0.0167]]),
  ('metric p-value sample 23',
   [[0.2, 0.2, 0.02, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.6, 0.6, 0.0835, 0.6, 0.0835]])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
all three pass their step thresholds[[True, True, True], [0.03, 0.03, 0.04]][[True, True, True], [0.03, 0.04, 0.03]]Failed
step-down stops at first non-rejection[[False, False, False], [0.06, 0.06, 0.06]][[False, False, False], [0.06, 0.06, 0.06]]Passed
boundary p equals alpha over remaining[[True, False, False, False], [0.05, 1.0, 1.0, 1.0]][[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]Passed
adjusted values are monotone[[True, True, False], [0.03, 0.03, 0.5]][[True, True, False], [0.03, 0.03, 0.5]]Passed
adjusted values cap at one[[False, False], [0.8, 0.8]][[False, False], [0.8, 0.8]]Passed
original order is preserved[[False, True, True], [0.003, 0.04, 0.3]][[False, True, True], [0.3, 0.003, 0.04]]Failed
metric p-value sample 1[[False, False, False, False], [0.08, 0.08, 0.1, 0.6]][[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]Passed
metric p-value sample 2[[True, True, True], [0.0375, 0.0375, 0.0375]][[True, True, True], [0.0375, 0.0375, 0.0375]]Passed

SHA-256 / f9ea11524a9bfc5380e140f401cafd1a04359a5507002d978d11721137f9194f

HELD IN THE MEMBER ARCHIVE

The verified repair and its recorded checks are member-only.

This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.

Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.

Member access is invitation-based. Sign in with your invited account to inspect the repair.

Sign in to the archive ↗

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.271047+00:00.

Case digest / f99ab756b28ca34fed853d8db01f93d57228e29651d9e0df93175dffeb727d15