FA-74641 / Experiment statistics / Open access
Guardrailed ship decision: A positive point estimate counts as a win · case 01
Launches ship on noise whenever the lift is positive.
ROOT CAUSE
The ship branch tests the point estimate primary[0] > 0.
VERIFIED REPAIR
Require the lower confidence bound to be above zero.
Unsuccessful approach: Accepting a lower bound equal to zero still ships inconclusive results.
Case contract
primary is [lift, ci_low, ci_high]. Each guardrail [name, ci_low, ci_high, margin, direction] passes non-inferiority when higher_is_better and ci_low > -margin, or lower_is_better and ci_high < margin. Decision: no_ship if primary ci_high < 0; else if primary ci_low > 0 then "blocked:<first failing guardrail in input order>" or ship; else inconclusive. Return [decision, failing names in input order].
Why this case matters
Launch reviews combine a primary win with guardrail checks; the combination rule must be exact.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > -margin
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[0] > 0:
if failing:
return ['blocked:' + failing[0], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 25', [[0.01, 0.0, 0.01], []], ['inconclusive', []]),
('readout sample 30',
[[0.01, 0.0, 0.01], [['latency', 0.004, 0.014, 0.02, 'lower_is_better']]],
['inconclusive', []])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 12',
[[0.02, 0.01, 0.04], [['latency', -0.03, -0.019999999999999997, 0.01, 'lower_is_better']]],
['ship', []])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']]),
('readout sample 25', [[0.01, 0.0, 0.01], []], ['inconclusive', []]),
('readout sample 32',
[[0.01, 0.0, 0.03],
[['latency', 0.0, 0.01, 0.02, 'higher_is_better'], ['revenue', 0.0, 0.01, 0.01, 'higher_is_better']]],
['inconclusive', []])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 22',
[[0.011, 0.001, 0.011], [['crashes', 0.004, 0.034, 0.01, 'higher_is_better']]],
['ship', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['ship', []] | ['ship', []] | Passed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['ship', []] | ['ship', []] | Passed |
| first failing guardrail is named | ['blocked:errors', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Passed |
| positive point estimate without significance | ['ship', []] | ['inconclusive', []] | Failed |
| interval touching zero is inconclusive | ['ship', []] | ['inconclusive', []] | Failed |
| readout sample 1 | ['blocked:latency', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Passed |
SHA-256 / c7c9c3d7b3ded10f5e58fe8771c5d2c50c5ea091416578e78ba3e46e682beae4
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > -margin
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[1] >= 0:
if failing:
return ['blocked:' + failing[0], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 25', [[0.01, 0.0, 0.01], []], ['inconclusive', []]),
('readout sample 30',
[[0.01, 0.0, 0.01], [['latency', 0.004, 0.014, 0.02, 'lower_is_better']]],
['inconclusive', []])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 12',
[[0.02, 0.01, 0.04], [['latency', -0.03, -0.019999999999999997, 0.01, 'lower_is_better']]],
['ship', []])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']]),
('readout sample 25', [[0.01, 0.0, 0.01], []], ['inconclusive', []]),
('readout sample 32',
[[0.01, 0.0, 0.03],
[['latency', 0.0, 0.01, 0.02, 'higher_is_better'], ['revenue', 0.0, 0.01, 0.01, 'higher_is_better']]],
['inconclusive', []])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 22',
[[0.011, 0.001, 0.011], [['crashes', 0.004, 0.034, 0.01, 'higher_is_better']]],
['ship', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['ship', []] | ['ship', []] | Passed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['ship', []] | ['ship', []] | Passed |
| first failing guardrail is named | ['blocked:errors', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Passed |
| positive point estimate without significance | ['inconclusive', []] | ['inconclusive', []] | Passed |
| interval touching zero is inconclusive | ['ship', []] | ['inconclusive', []] | Failed |
| readout sample 1 | ['blocked:latency', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Passed |
SHA-256 / 988bcc73d1d77135efe189eaa6c38e8a33b01c611d852a951c8409e416970a15
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > -margin
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[1] > 0:
if failing:
return ['blocked:' + failing[0], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 25', [[0.01, 0.0, 0.01], []], ['inconclusive', []]),
('readout sample 30',
[[0.01, 0.0, 0.01], [['latency', 0.004, 0.014, 0.02, 'lower_is_better']]],
['inconclusive', []])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 12',
[[0.02, 0.01, 0.04], [['latency', -0.03, -0.019999999999999997, 0.01, 'lower_is_better']]],
['ship', []])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']]),
('readout sample 25', [[0.01, 0.0, 0.01], []], ['inconclusive', []]),
('readout sample 32',
[[0.01, 0.0, 0.03],
[['latency', 0.0, 0.01, 0.02, 'higher_is_better'], ['revenue', 0.0, 0.01, 0.01, 'higher_is_better']]],
['inconclusive', []])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 22',
[[0.011, 0.001, 0.011], [['crashes', 0.004, 0.034, 0.01, 'higher_is_better']]],
['ship', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['ship', []] | ['ship', []] | Passed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['ship', []] | ['ship', []] | Passed |
| first failing guardrail is named | ['blocked:errors', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Passed |
| positive point estimate without significance | ['inconclusive', []] | ['inconclusive', []] | Passed |
| interval touching zero is inconclusive | ['inconclusive', []] | ['inconclusive', []] | Passed |
| readout sample 1 | ['blocked:latency', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Passed |
SHA-256 / 3a30f247ee1178fff72781d5341408c944386fd04864775bf4b24ef38927aa8b
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.619939+00:00.
Case digest / a06b880b337d946cb63a1dfbec07cbce8b031360921ce42a3e1e06edb05835d4