FAILURE MAP
← Case archive

FA-74626 / Experiment statistics / Open access

Guardrailed ship decision: Non-inferiority margin has the wrong sign · case 01

Higher-is-better guardrails demand a gain larger than the margin.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The check is ci_low > margin instead of ci_low > -margin.

VERIFIED REPAIR

Allow losses down to -margin.

Unsuccessful approach: Requiring ci_low > 0 turns non-inferiority into superiority.

Case contract

primary is [lift, ci_low, ci_high]. Each guardrail [name, ci_low, ci_high, margin, direction] passes non-inferiority when higher_is_better and ci_low > -margin, or lower_is_better and ci_high < margin. Decision: no_ship if primary ci_high < 0; else if primary ci_low > 0 then "blocked:<first failing guardrail in input order>" or ship; else inconclusive. Return [decision, failing names in input order].

Why this case matters

Launch reviews combine a primary win with guardrail checks; the combination rule must be exact.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(primary, guardrails):
    failing = []
    for name, lo, hi, margin, direction in guardrails:
        if direction == 'higher_is_better':
            ok = lo > margin
        else:
            ok = hi < margin
        if not ok:
            failing.append(name)
    if primary[2] < 0:
        return ['no_ship', failing]
    if primary[1] > 0:
        if failing:
            return ['blocked:' + failing[0], failing]
        return ['ship', failing]
    return ['inconclusive', failing]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
   [[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
   ['blocked:latency', ['latency']]),
  ('non-inferiority allows a small loss',
   [[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('readout sample 1',
   [[0.02, 0.01, 0.04],
    [['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
     ['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
   ['blocked:latency', ['latency', 'crashes']]),
  ('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
  ('readout sample 3',
   [[0.01, 0.0, 0.01],
    [['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
     ['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
   ['inconclusive', ['latency']])],
 [('non-inferiority allows a small loss',
   [[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('readout sample 6',
   [[0.01, 0.0, 0.01],
    [['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
     ['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
     ['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
   ['inconclusive', ['errors']]),
  ('readout sample 11',
   [[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('readout sample 13',
   [[-0.04, -0.05, -0.04],
    [['latency', -0.03, 0.0, 0.02, 'lower_is_better'],
     ['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
     ['crashes', -0.01, 0.0, 0.02, 'higher_is_better']]],
   ['no_ship', []])],
 [('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('readout sample 11',
   [[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('readout sample 28',
   [[0.02, 0.01, 0.04],
    [['revenue', 0.004, 0.054000000000000006, 0.02, 'higher_is_better'],
     ['crashes', 0.0, 0.01, 0.02, 'higher_is_better'],
     ['latency', -0.03, 0.0, 0.02, 'higher_is_better']]],
   ['blocked:latency', ['latency']]),
  ('readout sample 57',
   [[0.02, 0.01, 0.02], [['crashes', -0.01, 0.019999999999999997, 0.02, 'higher_is_better']]],
   ['ship', []])],
 [('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('significant loss is no ship',
   [[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
   ['no_ship', ['latency']]),
  ('readout sample 9',
   [[0.02, 0.01, 0.02],
    [['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
     ['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
     ['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'latency']]),
  ('readout sample 16',
   [[-0.01, -0.02, 0.009999999999999998],
    [['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
     ['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
     ['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
   ['inconclusive', ['errors', 'crashes']]),
  ('readout sample 58',
   [[0.02, 0.01, 0.02], [['latency', 0.0, 0.01, 0.01, 'higher_is_better']]],
   ['ship', []])],
 [('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('significant loss is no ship',
   [[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
   ['no_ship', ['latency']]),
  ('guardrail loss beyond margin',
   [[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('readout sample 9',
   [[0.02, 0.01, 0.02],
    [['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
     ['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
     ['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'latency']]),
  ('readout sample 21',
   [[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
   ['inconclusive', ['latency']]),
  ('readout sample 43',
   [[0.01, 0.0, 0.03], [['errors', 0.0, 0.01, 0.02, 'higher_is_better']]],
   ['inconclusive', []])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
latency guardrail uses the upper bound['blocked:latency', ['latency']]['blocked:latency', ['latency']]Passed
non-inferiority allows a small loss['blocked:revenue', ['revenue']]['ship', []]Failed
bound exactly at the margin fails['blocked:revenue', ['revenue']]['blocked:revenue', ['revenue']]Passed
bound just inside the margin passes['blocked:revenue', ['revenue']]['ship', []]Failed
first failing guardrail is named['blocked:errors', ['errors', 'crashes']]['blocked:errors', ['errors', 'crashes']]Passed
readout sample 1['blocked:latency', ['latency', 'crashes']]['blocked:latency', ['latency', 'crashes']]Passed
readout sample 2['no_ship', []]['no_ship', []]Passed
readout sample 3['inconclusive', ['revenue', 'latency']]['inconclusive', ['latency']]Failed

SHA-256 / c942fd632c4b8abb33162f5a18f8099ec4f60ddfe48e9be7576cbf9327f4d9a5

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(primary, guardrails):
    failing = []
    for name, lo, hi, margin, direction in guardrails:
        if direction == 'higher_is_better':
            ok = lo > 0
        else:
            ok = hi < margin
        if not ok:
            failing.append(name)
    if primary[2] < 0:
        return ['no_ship', failing]
    if primary[1] > 0:
        if failing:
            return ['blocked:' + failing[0], failing]
        return ['ship', failing]
    return ['inconclusive', failing]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
   [[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
   ['blocked:latency', ['latency']]),
  ('non-inferiority allows a small loss',
   [[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('readout sample 1',
   [[0.02, 0.01, 0.04],
    [['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
     ['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
   ['blocked:latency', ['latency', 'crashes']]),
  ('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
  ('readout sample 3',
   [[0.01, 0.0, 0.01],
    [['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
     ['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
   ['inconclusive', ['latency']])],
 [('non-inferiority allows a small loss',
   [[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('readout sample 6',
   [[0.01, 0.0, 0.01],
    [['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
     ['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
     ['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
   ['inconclusive', ['errors']]),
  ('readout sample 11',
   [[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('readout sample 13',
   [[-0.04, -0.05, -0.04],
    [['latency', -0.03, 0.0, 0.02, 'lower_is_better'],
     ['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
     ['crashes', -0.01, 0.0, 0.02, 'higher_is_better']]],
   ['no_ship', []])],
 [('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('readout sample 11',
   [[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('readout sample 28',
   [[0.02, 0.01, 0.04],
    [['revenue', 0.004, 0.054000000000000006, 0.02, 'higher_is_better'],
     ['crashes', 0.0, 0.01, 0.02, 'higher_is_better'],
     ['latency', -0.03, 0.0, 0.02, 'higher_is_better']]],
   ['blocked:latency', ['latency']]),
  ('readout sample 57',
   [[0.02, 0.01, 0.02], [['crashes', -0.01, 0.019999999999999997, 0.02, 'higher_is_better']]],
   ['ship', []])],
 [('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('significant loss is no ship',
   [[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
   ['no_ship', ['latency']]),
  ('readout sample 9',
   [[0.02, 0.01, 0.02],
    [['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
     ['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
     ['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'latency']]),
  ('readout sample 16',
   [[-0.01, -0.02, 0.009999999999999998],
    [['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
     ['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
     ['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
   ['inconclusive', ['errors', 'crashes']]),
  ('readout sample 58',
   [[0.02, 0.01, 0.02], [['latency', 0.0, 0.01, 0.01, 'higher_is_better']]],
   ['ship', []])],
 [('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('significant loss is no ship',
   [[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
   ['no_ship', ['latency']]),
  ('guardrail loss beyond margin',
   [[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('readout sample 9',
   [[0.02, 0.01, 0.02],
    [['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
     ['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
     ['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'latency']]),
  ('readout sample 21',
   [[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
   ['inconclusive', ['latency']]),
  ('readout sample 43',
   [[0.01, 0.0, 0.03], [['errors', 0.0, 0.01, 0.02, 'higher_is_better']]],
   ['inconclusive', []])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
latency guardrail uses the upper bound['blocked:latency', ['latency']]['blocked:latency', ['latency']]Passed
non-inferiority allows a small loss['blocked:revenue', ['revenue']]['ship', []]Failed
bound exactly at the margin fails['blocked:revenue', ['revenue']]['blocked:revenue', ['revenue']]Passed
bound just inside the margin passes['blocked:revenue', ['revenue']]['ship', []]Failed
first failing guardrail is named['blocked:errors', ['errors', 'crashes']]['blocked:errors', ['errors', 'crashes']]Passed
readout sample 1['blocked:latency', ['latency', 'crashes']]['blocked:latency', ['latency', 'crashes']]Passed
readout sample 2['no_ship', []]['no_ship', []]Passed
readout sample 3['inconclusive', ['revenue', 'latency']]['inconclusive', ['latency']]Failed

SHA-256 / 964f344fc0e18c7715e49b57a0f7ae5e74cea4ab82d7a43677b6c3c2701a34bc

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(primary, guardrails):
    failing = []
    for name, lo, hi, margin, direction in guardrails:
        if direction == 'higher_is_better':
            ok = lo > -margin
        else:
            ok = hi < margin
        if not ok:
            failing.append(name)
    if primary[2] < 0:
        return ['no_ship', failing]
    if primary[1] > 0:
        if failing:
            return ['blocked:' + failing[0], failing]
        return ['ship', failing]
    return ['inconclusive', failing]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('latency guardrail uses the upper bound',
   [[0.02, 0.01, 0.03], [['latency', -0.01, 0.025, 0.02, 'lower_is_better']]],
   ['blocked:latency', ['latency']]),
  ('non-inferiority allows a small loss',
   [[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('readout sample 1',
   [[0.02, 0.01, 0.04],
    [['latency', -0.03, 0.0, 0.01, 'higher_is_better'],
     ['crashes', -0.03, 0.020000000000000004, 0.01, 'lower_is_better']]],
   ['blocked:latency', ['latency', 'crashes']]),
  ('readout sample 2', [[-0.04, -0.05, -0.020000000000000004], []], ['no_ship', []]),
  ('readout sample 3',
   [[0.01, 0.0, 0.01],
    [['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
     ['latency', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
   ['inconclusive', ['latency']])],
 [('non-inferiority allows a small loss',
   [[0.02, 0.01, 0.03], [['revenue', -0.015, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('readout sample 6',
   [[0.01, 0.0, 0.01],
    [['errors', -0.03, 0.0, 0.02, 'higher_is_better'],
     ['crashes', 0.004, 0.014, 0.02, 'lower_is_better'],
     ['revenue', -0.01, 0.019999999999999997, 0.02, 'lower_is_better']]],
   ['inconclusive', ['errors']]),
  ('readout sample 11',
   [[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('readout sample 13',
   [[-0.04, -0.05, -0.04],
    [['latency', -0.03, 0.0, 0.02, 'lower_is_better'],
     ['revenue', 0.0, 0.03, 0.01, 'higher_is_better'],
     ['crashes', -0.01, 0.0, 0.02, 'higher_is_better']]],
   ['no_ship', []])],
 [('bound exactly at the margin fails',
   [[0.02, 0.01, 0.03], [['revenue', -0.02, 0.01, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('readout sample 11',
   [[0.02, 0.01, 0.02], [['revenue', -0.01, 0.04, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('readout sample 28',
   [[0.02, 0.01, 0.04],
    [['revenue', 0.004, 0.054000000000000006, 0.02, 'higher_is_better'],
     ['crashes', 0.0, 0.01, 0.02, 'higher_is_better'],
     ['latency', -0.03, 0.0, 0.02, 'higher_is_better']]],
   ['blocked:latency', ['latency']]),
  ('readout sample 57',
   [[0.02, 0.01, 0.02], [['crashes', -0.01, 0.019999999999999997, 0.02, 'higher_is_better']]],
   ['ship', []])],
 [('bound just inside the margin passes',
   [[0.02, 0.01, 0.03], [['revenue', -0.0196, 0.01, 0.02, 'higher_is_better']]],
   ['ship', []]),
  ('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('significant loss is no ship',
   [[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
   ['no_ship', ['latency']]),
  ('readout sample 9',
   [[0.02, 0.01, 0.02],
    [['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
     ['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
     ['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'latency']]),
  ('readout sample 16',
   [[-0.01, -0.02, 0.009999999999999998],
    [['latency', 0.004, 0.014, 0.02, 'higher_is_better'],
     ['errors', -0.03, 0.020000000000000004, 0.01, 'lower_is_better'],
     ['crashes', 0.004, 0.054000000000000006, 0.02, 'lower_is_better']]],
   ['inconclusive', ['errors', 'crashes']]),
  ('readout sample 58',
   [[0.02, 0.01, 0.02], [['latency', 0.0, 0.01, 0.01, 'higher_is_better']]],
   ['ship', []])],
 [('first failing guardrail is named',
   [[0.02, 0.01, 0.03],
    [['errors', 0.0, 0.05, 0.02, 'lower_is_better'], ['crashes', 0.0, 0.03, 0.01, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'crashes']]),
  ('positive point estimate without significance', [[0.01, -0.005, 0.025], []], ['inconclusive', []]),
  ('interval touching zero is inconclusive', [[0.01, 0.0, 0.02], []], ['inconclusive', []]),
  ('significant loss is no ship',
   [[-0.02, -0.03, -0.01], [['latency', 0.0, 0.1, 0.02, 'lower_is_better']]],
   ['no_ship', ['latency']]),
  ('guardrail loss beyond margin',
   [[0.02, 0.01, 0.03], [['revenue', -0.03, 0.0, 0.02, 'higher_is_better']]],
   ['blocked:revenue', ['revenue']]),
  ('readout sample 9',
   [[0.02, 0.01, 0.02],
    [['revenue', 0.0, 0.05, 0.01, 'higher_is_better'],
     ['errors', 0.0, 0.03, 0.01, 'lower_is_better'],
     ['latency', 0.0, 0.03, 0.02, 'lower_is_better']]],
   ['blocked:errors', ['errors', 'latency']]),
  ('readout sample 21',
   [[0.01, 0.0, 0.03], [['latency', -0.02, 0.030000000000000002, 0.01, 'lower_is_better']]],
   ['inconclusive', ['latency']]),
  ('readout sample 43',
   [[0.01, 0.0, 0.03], [['errors', 0.0, 0.01, 0.02, 'higher_is_better']]],
   ['inconclusive', []])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
latency guardrail uses the upper bound['blocked:latency', ['latency']]['blocked:latency', ['latency']]Passed
non-inferiority allows a small loss['ship', []]['ship', []]Passed
bound exactly at the margin fails['blocked:revenue', ['revenue']]['blocked:revenue', ['revenue']]Passed
bound just inside the margin passes['ship', []]['ship', []]Passed
first failing guardrail is named['blocked:errors', ['errors', 'crashes']]['blocked:errors', ['errors', 'crashes']]Passed
readout sample 1['blocked:latency', ['latency', 'crashes']]['blocked:latency', ['latency', 'crashes']]Passed
readout sample 2['no_ship', []]['no_ship', []]Passed
readout sample 3['inconclusive', ['latency']]['inconclusive', ['latency']]Passed

SHA-256 / af7f4941080df85c20ecead8923a476ab7f30a3a100fdbc66b1a086d5caa7e5c

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.522385+00:00.

Case digest / 2429ba8a1d5911b51deb25fcce32e4a58483830e21439ccf86b78a7aedb6f9db