FAILURE MAP
← Case archive

FA-74481 / Experiment statistics / Open access

Holm correction across metrics: Every step uses the Bonferroni threshold · case 01

Metrics that Holm would reject after the first step stay unrejected.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

Thresholds are alpha / m at every rank.

THE FAILURE

Thresholds are alpha / m at every rank.

Unsuccessful approach: The Benjamini-Hochberg threshold alpha (rank + 1) / m controls a different error rate.

Case contract

Holm step-down: sort p-values ascending (ties by position); reject while p_(k) <= alpha / (m - k) for k = 0, 1, ... and stop at the first failure. Adjusted p_(k) = running max of min(1, (m - k) p_(k)). Return [rejections, adjusted p-values rounded to 6] in the original metric order.

Why this case matters

Experiments track many metrics; multiplicity control keeps secondary wins honest.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(pvalues, alpha):
    m = len(pvalues)
    order = sorted(range(m), key=lambda i: (pvalues[i], i))
    reject = [False] * m
    adjusted = [0.0] * m
    running = 0.0
    stopped = False
    for rank, i in enumerate(order):
        running = max(running, min(1.0, (m - rank) * pvalues[i]))
        adjusted[i] = round(running, 6)
        if not stopped and pvalues[i] <= alpha / m:
            reject[i] = True
        else:
            stopped = True
    return [reject, adjusted]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 1',
   [[0.025, 0.02, 0.05, 0.6], 0.05],
   [[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]),
  ('metric p-value sample 12',
   [[0.04, 0.025, 0.01, 0.0167, 0.0125], 0.05],
   [[False, False, True, False, True], [0.0501, 0.0501, 0.05, 0.0501, 0.05]])],
 [('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 6', [[0.01, 0.05], 0.05], [[True, True], [0.02, 0.05]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 34',
   [[0.01, 0.0167, 0.001, 0.05, 0.04], 0.05],
   [[True, False, True, False, False], [0.04, 0.0501, 0.005, 0.08, 0.08]])],
 [('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 18',
   [[0.6, 0.001, 0.04, 0.025, 0.001], 0.05],
   [[False, True, False, False, True], [0.6, 0.005, 0.08, 0.075, 0.005]])],
 [('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 16',
   [[0.03, 0.02, 0.2, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.09, 0.0835, 0.4, 0.6, 0.0835]]),
  ('metric p-value sample 52',
   [[0.04, 0.01, 0.025, 0.2, 0.0167], 0.05],
   [[False, True, False, False, False], [0.08, 0.05, 0.075, 0.2, 0.0668]])],
 [('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 21',
   [[0.0167, 0.6, 0.2, 0.02, 0.01], 0.05],
   [[False, False, False, False, True], [0.0668, 0.6, 0.4, 0.0668, 0.05]]),
  ('metric p-value sample 22', [[0.0167], 0.05], [[True], [0.0167]]),
  ('metric p-value sample 23',
   [[0.2, 0.2, 0.02, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.6, 0.6, 0.0835, 0.6, 0.0835]])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
all three pass their step thresholds[[True, False, True], [0.03, 0.04, 0.03]][[True, True, True], [0.03, 0.04, 0.03]]Failed
step-down stops at first non-rejection[[False, False, False], [0.06, 0.06, 0.06]][[False, False, False], [0.06, 0.06, 0.06]]Passed
boundary p equals alpha over remaining[[True, False, False, False], [0.05, 1.0, 1.0, 1.0]][[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]Passed
adjusted values are monotone[[True, True, False], [0.03, 0.03, 0.5]][[True, True, False], [0.03, 0.03, 0.5]]Passed
adjusted values cap at one[[False, False], [0.8, 0.8]][[False, False], [0.8, 0.8]]Passed
original order is preserved[[False, True, False], [0.3, 0.003, 0.04]][[False, True, True], [0.3, 0.003, 0.04]]Failed
metric p-value sample 1[[False, False, False, False], [0.08, 0.08, 0.1, 0.6]][[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]Passed
metric p-value sample 12[[False, False, True, False, False], [0.0501, 0.0501, 0.05, 0.0501, 0.05]][[False, False, True, False, True], [0.0501, 0.0501, 0.05, 0.0501, 0.05]]Failed

SHA-256 / 77a2eb50bdecfd08ea3f6e3b99da8ec2857cc1a95341c223234987a731d6767e

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(pvalues, alpha):
    m = len(pvalues)
    order = sorted(range(m), key=lambda i: (pvalues[i], i))
    reject = [False] * m
    adjusted = [0.0] * m
    running = 0.0
    stopped = False
    for rank, i in enumerate(order):
        running = max(running, min(1.0, (m - rank) * pvalues[i]))
        adjusted[i] = round(running, 6)
        if not stopped and pvalues[i] <= alpha * (rank + 1) / m:
            reject[i] = True
        else:
            stopped = True
    return [reject, adjusted]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 1',
   [[0.025, 0.02, 0.05, 0.6], 0.05],
   [[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]),
  ('metric p-value sample 12',
   [[0.04, 0.025, 0.01, 0.0167, 0.0125], 0.05],
   [[False, False, True, False, True], [0.0501, 0.0501, 0.05, 0.0501, 0.05]])],
 [('step-down stops at first non-rejection',
   [[0.03, 0.02, 0.04], 0.05],
   [[False, False, False], [0.06, 0.06, 0.06]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('metric p-value sample 6', [[0.01, 0.05], 0.05], [[True, True], [0.02, 0.05]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 34',
   [[0.01, 0.0167, 0.001, 0.05, 0.04], 0.05],
   [[True, False, True, False, False], [0.04, 0.0501, 0.005, 0.08, 0.08]])],
 [('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('boundary p equals alpha over remaining',
   [[0.0125, 0.5, 0.9, 0.9], 0.05],
   [[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]),
  ('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 18',
   [[0.6, 0.001, 0.04, 0.025, 0.001], 0.05],
   [[False, True, False, False, True], [0.6, 0.005, 0.08, 0.075, 0.005]])],
 [('adjusted values are monotone', [[0.01, 0.011, 0.5], 0.05], [[True, True, False], [0.03, 0.03, 0.5]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 11', [[0.01, 0.001, 0.02], 0.05], [[True, True, True], [0.02, 0.003, 0.02]]),
  ('metric p-value sample 16',
   [[0.03, 0.02, 0.2, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.09, 0.0835, 0.4, 0.6, 0.0835]]),
  ('metric p-value sample 52',
   [[0.04, 0.01, 0.025, 0.2, 0.0167], 0.05],
   [[False, True, False, False, False], [0.08, 0.05, 0.075, 0.2, 0.0668]])],
 [('all three pass their step thresholds',
   [[0.01, 0.04, 0.012], 0.05],
   [[True, True, True], [0.03, 0.04, 0.03]]),
  ('adjusted values cap at one', [[0.4, 0.6], 0.05], [[False, False], [0.8, 0.8]]),
  ('original order is preserved', [[0.3, 0.001, 0.02], 0.05], [[False, True, True], [0.3, 0.003, 0.04]]),
  ('holm rejects more than bonferroni', [[0.01, 0.02], 0.05], [[True, True], [0.02, 0.02]]),
  ('single metric', [[0.05], 0.05], [[True], [0.05]]),
  ('metric p-value sample 21',
   [[0.0167, 0.6, 0.2, 0.02, 0.01], 0.05],
   [[False, False, False, False, True], [0.0668, 0.6, 0.4, 0.0668, 0.05]]),
  ('metric p-value sample 22', [[0.0167], 0.05], [[True], [0.0167]]),
  ('metric p-value sample 23',
   [[0.2, 0.2, 0.02, 0.6, 0.0167], 0.05],
   [[False, False, False, False, False], [0.6, 0.6, 0.0835, 0.6, 0.0835]])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
all three pass their step thresholds[[True, True, True], [0.03, 0.04, 0.03]][[True, True, True], [0.03, 0.04, 0.03]]Passed
step-down stops at first non-rejection[[False, False, False], [0.06, 0.06, 0.06]][[False, False, False], [0.06, 0.06, 0.06]]Passed
boundary p equals alpha over remaining[[True, False, False, False], [0.05, 1.0, 1.0, 1.0]][[True, False, False, False], [0.05, 1.0, 1.0, 1.0]]Passed
adjusted values are monotone[[True, True, False], [0.03, 0.03, 0.5]][[True, True, False], [0.03, 0.03, 0.5]]Passed
adjusted values cap at one[[False, False], [0.8, 0.8]][[False, False], [0.8, 0.8]]Passed
original order is preserved[[False, True, True], [0.3, 0.003, 0.04]][[False, True, True], [0.3, 0.003, 0.04]]Passed
metric p-value sample 1[[False, False, False, False], [0.08, 0.08, 0.1, 0.6]][[False, False, False, False], [0.08, 0.08, 0.1, 0.6]]Passed
metric p-value sample 12[[True, True, True, True, True], [0.0501, 0.0501, 0.05, 0.0501, 0.05]][[False, False, True, False, True], [0.0501, 0.0501, 0.05, 0.0501, 0.05]]Failed

SHA-256 / 46f780b052b7d5fd0c12a2b9b340092aac0bf9bb78155fbc16cc45cef98d22da

HELD IN THE MEMBER ARCHIVE

The verified repair and its recorded checks are member-only.

This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.

Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.

Member access is invitation-based. Sign in with your invited account to inspect the repair.

Sign in to the archive ↗

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:57.236427+00:00.

Case digest / 828b1c5f6aab067fa4d85e609101019f45400a5fcfdc2dfe5e1bcb44e4625631