FA-74206 / Feature flag rollout bucketing / Open access
Guarded progressive ramp: An error rate equal to the threshold rolls back · case 01
Launches sitting exactly at the tolerated 2 percent error rate are rolled back.
ROOT CAUSE
The breach test uses rate >= 0.02.
VERIFIED REPAIR
Roll back only when the error rate strictly exceeds 0.02.
Unsuccessful approach: Rounding the rate to two decimals first hides breaches such as 0.024.
Case contract
stages is an increasing list of percentages; the ramp starts at stages[0]. Each check [error_rate, sample] is processed in order: after a rollback everything stays at 0; a sample below 100 holds the current stage; an error rate above 0.02 rolls back to 0 permanently; otherwise advance one stage (staying at the last). State is complete at the last stage, rolled_back after a rollback, else ramping. Return [final percent, state, percent after each check].
Why this case matters
Automated ramps with guardrails must neither overreact to noise nor resume after a rollback.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(stages, checks):
idx = 0
history = []
state = 'ramping' if len(stages) > 1 else 'complete'
for rate, sample in checks:
if state == 'rolled_back':
history.append(0)
continue
if sample < 100:
history.append(stages[idx])
continue
if rate >= 0.02:
state = 'rolled_back'
history.append(0)
continue
idx = min(idx + 1, len(stages) - 1)
if idx == len(stages) - 1:
state = 'complete'
history.append(stages[idx])
final = 0 if state == 'rolled_back' else stages[idx]
return [final, state, history]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),
('guardrail sample 2',
[[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],
[('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('guardrail sample 6',
[[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],
[0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),
('guardrail sample 28',
[[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],
[0, 'rolled_back', [100, 100, 100, 0, 0]]),
('guardrail sample 31',
[[100], [[0.05, 99], [0.01, 50], [0.02, 500], [0.02, 100], [0.01, 500], [0.05, 500]]],
[0, 'rolled_back', [100, 100, 100, 100, 100, 0]])],
[('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('guardrail sample 11',
[[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),
('guardrail sample 51',
[[1, 5, 25, 100], [[0.05, 99], [0.05, 99], [0.02, 100], [0.05, 99], [0.024, 500], [0.01, 100]]],
[0, 'rolled_back', [1, 1, 5, 5, 0, 0]])],
[('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 16',
[[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],
[0, 'rolled_back', [5, 5, 0, 0, 0]]),
('guardrail sample 18',
[[10, 50, 100], [[0.024, 500], [0.05, 50], [0.02, 100], [0.02, 500], [0.05, 500], [0.024, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 28',
[[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],
[0, 'rolled_back', [100, 100, 100, 0, 0]])],
[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 21',
[[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],
[100, 'complete', [5, 100, 100, 100, 100, 100]]),
('guardrail sample 22',
[[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],
[0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),
('guardrail sample 40',
[[1, 5, 25, 100], [[0.02, 100], [0.01, 500], [0.024, 100], [0, 99], [0.02, 500], [0.01, 500]]],
[0, 'rolled_back', [5, 25, 0, 0, 0, 0]])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| low-sample breach only holds | [5, 'ramping', [1, 5]] | [5, 'ramping', [1, 5]] | Passed |
| rollback is terminal | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| error rate exactly at threshold advances | [0, 'rolled_back', [0]] | [50, 'ramping', [50]] | Failed |
| error rate just above threshold rolls back | [0, 'rolled_back', [0]] | [0, 'rolled_back', [0]] | Passed |
| history records the stage after advancing | [25, 'ramping', [5, 25]] | [25, 'ramping', [5, 25]] | Passed |
| guardrail sample 1 | [1, 'ramping', [1, 1, 1]] | [1, 'ramping', [1, 1, 1]] | Passed |
| guardrail sample 2 | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| guardrail sample 3 | [0, 'rolled_back', [1, 0]] | [0, 'rolled_back', [1, 0]] | Passed |
SHA-256 / 9982017b58f7ba6132892684c60844431724db3f303a3e4c963495b831342eef
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(stages, checks):
idx = 0
history = []
state = 'ramping' if len(stages) > 1 else 'complete'
for rate, sample in checks:
if state == 'rolled_back':
history.append(0)
continue
if sample < 100:
history.append(stages[idx])
continue
if round(rate, 2) > 0.02:
state = 'rolled_back'
history.append(0)
continue
idx = min(idx + 1, len(stages) - 1)
if idx == len(stages) - 1:
state = 'complete'
history.append(stages[idx])
final = 0 if state == 'rolled_back' else stages[idx]
return [final, state, history]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),
('guardrail sample 2',
[[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],
[('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('guardrail sample 6',
[[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],
[0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),
('guardrail sample 28',
[[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],
[0, 'rolled_back', [100, 100, 100, 0, 0]]),
('guardrail sample 31',
[[100], [[0.05, 99], [0.01, 50], [0.02, 500], [0.02, 100], [0.01, 500], [0.05, 500]]],
[0, 'rolled_back', [100, 100, 100, 100, 100, 0]])],
[('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('guardrail sample 11',
[[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),
('guardrail sample 51',
[[1, 5, 25, 100], [[0.05, 99], [0.05, 99], [0.02, 100], [0.05, 99], [0.024, 500], [0.01, 100]]],
[0, 'rolled_back', [1, 1, 5, 5, 0, 0]])],
[('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 16',
[[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],
[0, 'rolled_back', [5, 5, 0, 0, 0]]),
('guardrail sample 18',
[[10, 50, 100], [[0.024, 500], [0.05, 50], [0.02, 100], [0.02, 500], [0.05, 500], [0.024, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 28',
[[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],
[0, 'rolled_back', [100, 100, 100, 0, 0]])],
[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 21',
[[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],
[100, 'complete', [5, 100, 100, 100, 100, 100]]),
('guardrail sample 22',
[[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],
[0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),
('guardrail sample 40',
[[1, 5, 25, 100], [[0.02, 100], [0.01, 500], [0.024, 100], [0, 99], [0.02, 500], [0.01, 500]]],
[0, 'rolled_back', [5, 25, 0, 0, 0, 0]])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| low-sample breach only holds | [5, 'ramping', [1, 5]] | [5, 'ramping', [1, 5]] | Passed |
| rollback is terminal | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| error rate exactly at threshold advances | [50, 'ramping', [50]] | [50, 'ramping', [50]] | Passed |
| error rate just above threshold rolls back | [50, 'ramping', [50]] | [0, 'rolled_back', [0]] | Failed |
| history records the stage after advancing | [25, 'ramping', [5, 25]] | [25, 'ramping', [5, 25]] | Passed |
| guardrail sample 1 | [1, 'ramping', [1, 1, 1]] | [1, 'ramping', [1, 1, 1]] | Passed |
| guardrail sample 2 | [0, 'rolled_back', [5, 5, 0]] | [0, 'rolled_back', [0, 0, 0]] | Failed |
| guardrail sample 3 | [0, 'rolled_back', [1, 0]] | [0, 'rolled_back', [1, 0]] | Passed |
SHA-256 / 318f0796ef2933b4742ad56692df980fc7faca787937d54e1b383a74968db0bc
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(stages, checks):
idx = 0
history = []
state = 'ramping' if len(stages) > 1 else 'complete'
for rate, sample in checks:
if state == 'rolled_back':
history.append(0)
continue
if sample < 100:
history.append(stages[idx])
continue
if rate > 0.02:
state = 'rolled_back'
history.append(0)
continue
idx = min(idx + 1, len(stages) - 1)
if idx == len(stages) - 1:
state = 'complete'
history.append(stages[idx])
final = 0 if state == 'rolled_back' else stages[idx]
return [final, state, history]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),
('guardrail sample 2',
[[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],
[('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('guardrail sample 6',
[[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],
[0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),
('guardrail sample 28',
[[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],
[0, 'rolled_back', [100, 100, 100, 0, 0]]),
('guardrail sample 31',
[[100], [[0.05, 99], [0.01, 50], [0.02, 500], [0.02, 100], [0.01, 500], [0.05, 500]]],
[0, 'rolled_back', [100, 100, 100, 100, 100, 0]])],
[('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('guardrail sample 11',
[[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),
('guardrail sample 51',
[[1, 5, 25, 100], [[0.05, 99], [0.05, 99], [0.02, 100], [0.05, 99], [0.024, 500], [0.01, 100]]],
[0, 'rolled_back', [1, 1, 5, 5, 0, 0]])],
[('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 16',
[[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],
[0, 'rolled_back', [5, 5, 0, 0, 0]]),
('guardrail sample 18',
[[10, 50, 100], [[0.024, 500], [0.05, 50], [0.02, 100], [0.02, 500], [0.05, 500], [0.024, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 28',
[[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],
[0, 'rolled_back', [100, 100, 100, 0, 0]])],
[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 21',
[[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],
[100, 'complete', [5, 100, 100, 100, 100, 100]]),
('guardrail sample 22',
[[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],
[0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),
('guardrail sample 40',
[[1, 5, 25, 100], [[0.02, 100], [0.01, 500], [0.024, 100], [0, 99], [0.02, 500], [0.01, 500]]],
[0, 'rolled_back', [5, 25, 0, 0, 0, 0]])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| low-sample breach only holds | [5, 'ramping', [1, 5]] | [5, 'ramping', [1, 5]] | Passed |
| rollback is terminal | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| error rate exactly at threshold advances | [50, 'ramping', [50]] | [50, 'ramping', [50]] | Passed |
| error rate just above threshold rolls back | [0, 'rolled_back', [0]] | [0, 'rolled_back', [0]] | Passed |
| history records the stage after advancing | [25, 'ramping', [5, 25]] | [25, 'ramping', [5, 25]] | Passed |
| guardrail sample 1 | [1, 'ramping', [1, 1, 1]] | [1, 'ramping', [1, 1, 1]] | Passed |
| guardrail sample 2 | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| guardrail sample 3 | [0, 'rolled_back', [1, 0]] | [0, 'rolled_back', [1, 0]] | Passed |
SHA-256 / 9b06d174c276b60a1221ce9af36e243f9dd75698de550e6ffee07a39526dff46
Verification & scope
A deterministic toy flag-evaluation model with a stipulated contract; it does not reproduce any vendor SDK byte for byte. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:54.559669+00:00.
Case digest / a761221dac5e0d98aaa0df648ff9f6aeae7929d76bb0bbcc1b4db5bb1245c755