FA-74626 / Experiment statistics / Open access
Guardrailed ship decision: Non-inferiority margin has the wrong sign · case 01
Higher-is-better guardrails demand a gain larger than the margin.
ROOT CAUSE
The check is ci_low > margin instead of ci_low > -margin.
VERIFIED REPAIR
Allow losses down to -margin.
Unsuccessful approach: Requiring ci_low > 0 turns non-inferiority into superiority.
Case contract
primary is [lift, ci_low, ci_high]. Each guardrail [name, ci_low, ci_high, margin, direction] passes non-inferiority when higher_is_better and ci_low > -margin, or lower_is_better and ci_high < margin. Decision: no_ship if primary ci_high < 0; else if primary ci_low > 0 then "blocked:<first failing guardrail in input order>" or ship; else inconclusive. Return [decision, failing names in input order].
Why this case matters
Launch reviews combine a primary win with guardrail checks; the combination rule must be exact.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > margin
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[1] > 0:
if failing:
return ['blocked:' + failing[0], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 13',
[[-0.04, -0.05, -0.04],
[['latency', -0.03, 0.0, 0.02, 'lower_is_better'],
['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['crashes', -0.01, 0.0, 0.02, 'higher_is_better']]],
['no_ship', []])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 28',
[[0.02, 0.01, 0.04],
[['revenue', 0.004, 0.054000000000000006, 0.02, 'higher_is_better'],
['crashes', 0.0, 0.01, 0.02, 'higher_is_better'],
['latency', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:latency', ['latency']]),
('readout sample 57',
[[0.02, 0.01, 0.02], [['crashes', -0.01, 0.019999999999999997, 0.02, 'higher_is_better']]],
['ship', []])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 9',
[[0.02, 0.01, 0.02],
[['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
['blocked:errors', ['errors', 'latency']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']]),
('readout sample 58',
[[0.02, 0.01, 0.02], [['latency', 0.0, 0.01, 0.01, 'higher_is_better']]],
['ship', []])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 9',
[[0.02, 0.01, 0.02],
[['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
['blocked:errors', ['errors', 'latency']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 43',
[[0.01, 0.0, 0.03], [['errors', 0.0, 0.01, 0.02, 'higher_is_better']]],
['inconclusive', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['blocked:revenue', ['revenue']] | ['ship', []] | Failed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['blocked:revenue', ['revenue']] | ['ship', []] | Failed |
| first failing guardrail is named | ['blocked:errors', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Passed |
| readout sample 1 | ['blocked:latency', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Passed |
| readout sample 2 | ['no_ship', []] | ['no_ship', []] | Passed |
| readout sample 3 | ['inconclusive', ['revenue', 'latency']] | ['inconclusive', ['latency']] | Failed |
SHA-256 / c942fd632c4b8abb33162f5a18f8099ec4f60ddfe48e9be7576cbf9327f4d9a5
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > 0
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[1] > 0:
if failing:
return ['blocked:' + failing[0], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 13',
[[-0.04, -0.05, -0.04],
[['latency', -0.03, 0.0, 0.02, 'lower_is_better'],
['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['crashes', -0.01, 0.0, 0.02, 'higher_is_better']]],
['no_ship', []])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 28',
[[0.02, 0.01, 0.04],
[['revenue', 0.004, 0.054000000000000006, 0.02, 'higher_is_better'],
['crashes', 0.0, 0.01, 0.02, 'higher_is_better'],
['latency', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:latency', ['latency']]),
('readout sample 57',
[[0.02, 0.01, 0.02], [['crashes', -0.01, 0.019999999999999997, 0.02, 'higher_is_better']]],
['ship', []])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 9',
[[0.02, 0.01, 0.02],
[['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
['blocked:errors', ['errors', 'latency']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']]),
('readout sample 58',
[[0.02, 0.01, 0.02], [['latency', 0.0, 0.01, 0.01, 'higher_is_better']]],
['ship', []])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 9',
[[0.02, 0.01, 0.02],
[['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
['blocked:errors', ['errors', 'latency']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 43',
[[0.01, 0.0, 0.03], [['errors', 0.0, 0.01, 0.02, 'higher_is_better']]],
['inconclusive', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['blocked:revenue', ['revenue']] | ['ship', []] | Failed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['blocked:revenue', ['revenue']] | ['ship', []] | Failed |
| first failing guardrail is named | ['blocked:errors', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Passed |
| readout sample 1 | ['blocked:latency', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Passed |
| readout sample 2 | ['no_ship', []] | ['no_ship', []] | Passed |
| readout sample 3 | ['inconclusive', ['revenue', 'latency']] | ['inconclusive', ['latency']] | Failed |
SHA-256 / 964f344fc0e18c7715e49b57a0f7ae5e74cea4ab82d7a43677b6c3c2701a34bc
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(primary, guardrails):
failing = []
for name, lo, hi, margin, direction in guardrails:
if direction == 'higher_is_better':
ok = lo > -margin
else:
ok = hi < margin
if not ok:
failing.append(name)
if primary[2] < 0:
return ['no_ship', failing]
if primary[1] > 0:
if failing:
return ['blocked:' + failing[0], failing]
return ['ship', failing]
return ['inconclusive', failing]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
[[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
['blocked:latency', ['latency']]),
('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('readout sample 1',
[[0.02, 0.01, 0.04],
[['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
['blocked:latency', ['latency', 'crashes']]),
('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
('readout sample 3',
[[0.01, 0.0, 0.01],
[['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['latency']])],
[('non-inferiority allows a small loss',
[[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('readout sample 6',
[[0.01, 0.0, 0.01],
[['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
['inconclusive', ['errors']]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 13',
[[-0.04, -0.05, -0.04],
[['latency', -0.03, 0.0, 0.02, 'lower_is_better'],
['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
['crashes', -0.01, 0.0, 0.02, 'higher_is_better']]],
['no_ship', []])],
[('bound exactly at the margin fails',
[[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('readout sample 11',
[[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
['ship', []]),
('readout sample 28',
[[0.02, 0.01, 0.04],
[['revenue', 0.004, 0.054000000000000006, 0.02, 'higher_is_better'],
['crashes', 0.0, 0.01, 0.02, 'higher_is_better'],
['latency', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:latency', ['latency']]),
('readout sample 57',
[[0.02, 0.01, 0.02], [['crashes', -0.01, 0.019999999999999997, 0.02, 'higher_is_better']]],
['ship', []])],
[('bound just inside the margin passes',
[[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
['ship', []]),
('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('readout sample 9',
[[0.02, 0.01, 0.02],
[['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
['blocked:errors', ['errors', 'latency']]),
('readout sample 16',
[[-0.01, -0.02, 0.009999999999999998],
[['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
['inconclusive', ['errors', 'crashes']]),
('readout sample 58',
[[0.02, 0.01, 0.02], [['latency', 0.0, 0.01, 0.01, 'higher_is_better']]],
['ship', []])],
[('first failing guardrail is named',
[[0.02, 0.01, 0.03],
[['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
['blocked:errors', ['errors', 'crashes']]),
('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
('significant loss is no ship',
[[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
['no_ship', ['latency']]),
('guardrail loss beyond margin',
[[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
['blocked:revenue', ['revenue']]),
('readout sample 9',
[[0.02, 0.01, 0.02],
[['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
['blocked:errors', ['errors', 'latency']]),
('readout sample 21',
[[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
['inconclusive', ['latency']]),
('readout sample 43',
[[0.01, 0.0, 0.03], [['errors', 0.0, 0.01, 0.02, 'higher_is_better']]],
['inconclusive', []])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| latency guardrail uses the upper bound | ['blocked:latency', ['latency']] | ['blocked:latency', ['latency']] | Passed |
| non-inferiority allows a small loss | ['ship', []] | ['ship', []] | Passed |
| bound exactly at the margin fails | ['blocked:revenue', ['revenue']] | ['blocked:revenue', ['revenue']] | Passed |
| bound just inside the margin passes | ['ship', []] | ['ship', []] | Passed |
| first failing guardrail is named | ['blocked:errors', ['errors', 'crashes']] | ['blocked:errors', ['errors', 'crashes']] | Passed |
| readout sample 1 | ['blocked:latency', ['latency', 'crashes']] | ['blocked:latency', ['latency', 'crashes']] | Passed |
| readout sample 2 | ['no_ship', []] | ['no_ship', []] | Passed |
| readout sample 3 | ['inconclusive', ['latency']] | ['inconclusive', ['latency']] | Passed |
SHA-256 / af7f4941080df85c20ecead8923a476ab7f30a3a100fdbc66b1a086d5caa7e5c
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.522385+00:00.
Case digest / 2429ba8a1d5911b51deb25fcce32e4a58483830e21439ccf86b78a7aedb6f9db