FA-74636 / Experiment statistics / Open access
Guardrailed ship decision: The last failing guardrail is reported · case 01
Reviews chase a minor guardrail while the first listed blocker is hidden.
ROOT CAUSE
The blocked decision names failing[-1].
THE FAILURE
The blocked decision names failing[-1].
Unsuccessful approach: Sorting names alphabetically still ignores the configured order.
Case contract
primary is [lift, ci_low, ci_high]. Each guardrail [name, ci_low, ci_high, margin, direction] passes non-inferiority when higher_is_better and ci_low > -margin, or lower_is_better and ci_high < margin. Decision: no_ship if primary ci_high < 0; else if primary ci_low > 0 then "blocked:<first failing guardrail in input order>" or ship; else inconclusive. Return [decision, failing names in input order].
Why this case matters
Launch reviews combine a primary win with guardrail checks; the combination rule must be exact.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > -margin
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[1] > 0:
if failing:
return ['blocked:' + failing[-1], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 50',
[[0.02, 0.01, 0.02],
[['latency', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', -0.02, -0.01, 0.01, 'lower_is_better'],
['revenue', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'revenue']]),
('readout sample 53',
[[0.011, 0.001, 0.031],
[['revenue', 0.0, 0.05, 0.01, 'lower_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['latency', -0.03, 0.0, 0.01, 'lower_is_better']]],
['blocked:revenue', ['revenue', 'errors']])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 31',
[[0.011, 0.001, 0.011],
[['crashes', -0.03, 0.0, 0.01, 'higher_is_better'],
['errors', 0.004, 0.014, 0.02, 'lower_is_better'],
['latency', -0.02, -0.01, 0.01, 'higher_is_better']]],
['blocked:crashes', ['crashes', 'latency']]),
('readout sample 35',
[[0.011, 0.001, 0.011],
[['errors', -0.03, 0.020000000000000004, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 14',
[[0.011, 0.001, 0.031],
[['crashes', -0.03, -0.019999999999999997, 0.01, 'higher_is_better'],
['revenue', 0.004, 0.034, 0.02, 'lower_is_better']]],
['blocked:crashes', ['crashes', 'revenue']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 22',
[[0.011, 0.001, 0.011], [['crashes', 0.004, 0.034, 0.01, 'higher_is_better']]],
['ship', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['ship', []] | ['ship', []] | Passed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['ship', []] | ['ship', []] | Passed |
| first failing guardrail is named | ['blocked:crashes', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Failed |
| readout sample 1 | ['blocked:crashes', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Failed |
| readout sample 2 | ['no_ship', []] | ['no_ship', []] | Passed |
| readout sample 3 | ['inconclusive', ['latency']] | ['inconclusive', ['latency']] | Passed |
SHA-256 / 513c671d064399a5ca323d6521b042cde84cebe9232eea29db1ddc5997a0ba8c
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > -margin
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[1] > 0:
if failing:
return ['blocked:' + sorted(failing)[0], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 50',
[[0.02, 0.01, 0.02],
[['latency', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', -0.02, -0.01, 0.01, 'lower_is_better'],
['revenue', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'revenue']]),
('readout sample 53',
[[0.011, 0.001, 0.031],
[['revenue', 0.0, 0.05, 0.01, 'lower_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['latency', -0.03, 0.0, 0.01, 'lower_is_better']]],
['blocked:revenue', ['revenue', 'errors']])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 31',
[[0.011, 0.001, 0.011],
[['crashes', -0.03, 0.0, 0.01, 'higher_is_better'],
['errors', 0.004, 0.014, 0.02, 'lower_is_better'],
['latency', -0.02, -0.01, 0.01, 'higher_is_better']]],
['blocked:crashes', ['crashes', 'latency']]),
('readout sample 35',
[[0.011, 0.001, 0.011],
[['errors', -0.03, 0.020000000000000004, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 14',
[[0.011, 0.001, 0.031],
[['crashes', -0.03, -0.019999999999999997, 0.01, 'higher_is_better'],
['revenue', 0.004, 0.034, 0.02, 'lower_is_better']]],
['blocked:crashes', ['crashes', 'revenue']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 22',
[[0.011, 0.001, 0.011], [['crashes', 0.004, 0.034, 0.01, 'higher_is_better']]],
['ship', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['ship', []] | ['ship', []] | Passed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['ship', []] | ['ship', []] | Passed |
| first failing guardrail is named | ['blocked:crashes', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Failed |
| readout sample 1 | ['blocked:crashes', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Failed |
| readout sample 2 | ['no_ship', []] | ['no_ship', []] | Passed |
| readout sample 3 | ['inconclusive', ['latency']] | ['inconclusive', ['latency']] | Passed |
SHA-256 / 5421bedb107971a9b7df0d5c44c206241cb8a7432c886d6b1bfb748c0f489f68
HELD IN THE MEMBER ARCHIVE
The verified repair and its recorded checks are member-only.
This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.
Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.
Member access is invitation-based. Sign in with your invited account to inspect the repair.
Sign in to the archive ↗Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.575381+00:00.
Case digest / af825984266e9f5a89c543ba64a609696d2f712a78e0d2e1ad55ff7c0e383d64