FA-74211 / Feature flag rollout bucketing / Open access
Guarded progressive ramp: History lags one stage behind · case 01
Audit history shows the percentage before each advance, hiding when the ramp reached full traffic.
ROOT CAUSE
The stage is appended to history before idx advances.
VERIFIED REPAIR
Append the stage after advancing.
Unsuccessful approach: Appending the next stage (one ahead of the advance) double counts progress.
Case contract
stages is an increasing list of percentages; the ramp starts at stages[0]. Each check [error_rate, sample] is processed in order: after a rollback everything stays at 0; a sample below 100 holds the current stage; an error rate above 0.02 rolls back to 0 permanently; otherwise advance one stage (staying at the last). State is complete at the last stage, rolled_back after a rollback, else ramping. Return [final percent, state, percent after each check].
Why this case matters
Automated ramps with guardrails must neither overreact to noise nor resume after a rollback.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(stages, checks):
idx = 0
history = []
state = 'ramping' if len(stages) > 1 else 'complete'
for rate, sample in checks:
if state == 'rolled_back':
history.append(0)
continue
if sample < 100:
history.append(stages[idx])
continue
if rate > 0.02:
state = 'rolled_back'
history.append(0)
continue
history.append(stages[idx])
idx = min(idx + 1, len(stages) - 1)
if idx == len(stages) - 1:
state = 'complete'
final = 0 if state == 'rolled_back' else stages[idx]
return [final, state, history]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),
('guardrail sample 2',
[[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],
[0, 'rolled_back', [0, 0, 0]])],
[('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 20', [[1, 5, 25, 100], [[0, 100], [0, 100], [0, 50]]], [25, 'ramping', [5, 25, 25]]),
('guardrail sample 24',
[[1, 5, 25, 100], [[0.01, 50], [0.01, 100], [0.02, 500], [0.024, 99], [0.01, 99], [0, 50]]],
[25, 'ramping', [1, 5, 25, 25, 25, 25]])],
[('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('guardrail sample 11',
[[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 32',
[[10, 50, 100], [[0.01, 99], [0.01, 500], [0, 500], [0.02, 100], [0.024, 50], [0.02, 500]]],
[100, 'complete', [10, 50, 100, 100, 100, 100]]),
('guardrail sample 39',
[[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],
[0, 'rolled_back', [50, 0, 0, 0]])],
[('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 16',
[[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],
[0, 'rolled_back', [5, 5, 0, 0, 0]]),
('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),
('guardrail sample 59',
[[1, 5, 25, 100], [[0.02, 50], [0.01, 100], [0.024, 50], [0.024, 50], [0.05, 99], [0.05, 50]]],
[5, 'ramping', [1, 5, 5, 5, 5, 5]])],
[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 21',
[[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],
[100, 'complete', [5, 100, 100, 100, 100, 100]]),
('guardrail sample 23',
[[10, 50, 100], [[0.01, 500], [0.024, 99], [0, 50]]],
[50, 'ramping', [50, 50, 50]])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| low-sample breach only holds | [5, 'ramping', [1, 1]] | [5, 'ramping', [1, 5]] | Failed |
| rollback is terminal | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| error rate exactly at threshold advances | [50, 'ramping', [10]] | [50, 'ramping', [50]] | Failed |
| error rate just above threshold rolls back | [0, 'rolled_back', [0]] | [0, 'rolled_back', [0]] | Passed |
| history records the stage after advancing | [25, 'ramping', [1, 5]] | [25, 'ramping', [5, 25]] | Failed |
| rollback after completion | [0, 'rolled_back', [10, 50, 0]] | [0, 'rolled_back', [50, 100, 0]] | Failed |
| guardrail sample 1 | [1, 'ramping', [1, 1, 1]] | [1, 'ramping', [1, 1, 1]] | Passed |
| guardrail sample 2 | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
SHA-256 / 2663a132b8d0eb6918f160a131f0a3905a5746434b09ab5a714e24e413bce454
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(stages, checks):
idx = 0
history = []
state = 'ramping' if len(stages) > 1 else 'complete'
for rate, sample in checks:
if state == 'rolled_back':
history.append(0)
continue
if sample < 100:
history.append(stages[idx])
continue
if rate > 0.02:
state = 'rolled_back'
history.append(0)
continue
idx = min(idx + 1, len(stages) - 1)
if idx == len(stages) - 1:
state = 'complete'
history.append(stages[min(idx + 1, len(stages) - 1)])
final = 0 if state == 'rolled_back' else stages[idx]
return [final, state, history]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),
('guardrail sample 2',
[[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],
[0, 'rolled_back', [0, 0, 0]])],
[('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 20', [[1, 5, 25, 100], [[0, 100], [0, 100], [0, 50]]], [25, 'ramping', [5, 25, 25]]),
('guardrail sample 24',
[[1, 5, 25, 100], [[0.01, 50], [0.01, 100], [0.02, 500], [0.024, 99], [0.01, 99], [0, 50]]],
[25, 'ramping', [1, 5, 25, 25, 25, 25]])],
[('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('guardrail sample 11',
[[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 32',
[[10, 50, 100], [[0.01, 99], [0.01, 500], [0, 500], [0.02, 100], [0.024, 50], [0.02, 500]]],
[100, 'complete', [10, 50, 100, 100, 100, 100]]),
('guardrail sample 39',
[[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],
[0, 'rolled_back', [50, 0, 0, 0]])],
[('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 16',
[[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],
[0, 'rolled_back', [5, 5, 0, 0, 0]]),
('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),
('guardrail sample 59',
[[1, 5, 25, 100], [[0.02, 50], [0.01, 100], [0.024, 50], [0.024, 50], [0.05, 99], [0.05, 50]]],
[5, 'ramping', [1, 5, 5, 5, 5, 5]])],
[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 21',
[[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],
[100, 'complete', [5, 100, 100, 100, 100, 100]]),
('guardrail sample 23',
[[10, 50, 100], [[0.01, 500], [0.024, 99], [0, 50]]],
[50, 'ramping', [50, 50, 50]])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| low-sample breach only holds | [5, 'ramping', [1, 25]] | [5, 'ramping', [1, 5]] | Failed |
| rollback is terminal | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| error rate exactly at threshold advances | [50, 'ramping', [100]] | [50, 'ramping', [50]] | Failed |
| error rate just above threshold rolls back | [0, 'rolled_back', [0]] | [0, 'rolled_back', [0]] | Passed |
| history records the stage after advancing | [25, 'ramping', [25, 100]] | [25, 'ramping', [5, 25]] | Failed |
| rollback after completion | [0, 'rolled_back', [100, 100, 0]] | [0, 'rolled_back', [50, 100, 0]] | Failed |
| guardrail sample 1 | [1, 'ramping', [1, 1, 1]] | [1, 'ramping', [1, 1, 1]] | Passed |
| guardrail sample 2 | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
SHA-256 / b75c2cd3ad194ab25fd6361b37595d2a86f50eeabfd1b3c281ab87a04aadec1f
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(stages, checks):
idx = 0
history = []
state = 'ramping' if len(stages) > 1 else 'complete'
for rate, sample in checks:
if state == 'rolled_back':
history.append(0)
continue
if sample < 100:
history.append(stages[idx])
continue
if rate > 0.02:
state = 'rolled_back'
history.append(0)
continue
idx = min(idx + 1, len(stages) - 1)
if idx == len(stages) - 1:
state = 'complete'
history.append(stages[idx])
final = 0 if state == 'rolled_back' else stages[idx]
return [final, state, history]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),
('guardrail sample 2',
[[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],
[0, 'rolled_back', [0, 0, 0]])],
[('rollback is terminal',
[[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],
[0, 'rolled_back', [0, 0, 0]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 20', [[1, 5, 25, 100], [[0, 100], [0, 100], [0, 50]]], [25, 'ramping', [5, 25, 25]]),
('guardrail sample 24',
[[1, 5, 25, 100], [[0.01, 50], [0.01, 100], [0.02, 500], [0.024, 99], [0.01, 99], [0, 50]]],
[25, 'ramping', [1, 5, 25, 25, 25, 25]])],
[('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('guardrail sample 11',
[[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],
[0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),
('guardrail sample 32',
[[10, 50, 100], [[0.01, 99], [0.01, 500], [0, 500], [0.02, 100], [0.024, 50], [0.02, 500]]],
[100, 'complete', [10, 50, 100, 100, 100, 100]]),
('guardrail sample 39',
[[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],
[0, 'rolled_back', [50, 0, 0, 0]])],
[('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 16',
[[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],
[0, 'rolled_back', [5, 5, 0, 0, 0]]),
('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),
('guardrail sample 59',
[[1, 5, 25, 100], [[0.02, 50], [0.01, 100], [0.024, 50], [0.024, 50], [0.05, 99], [0.05, 50]]],
[5, 'ramping', [1, 5, 5, 5, 5, 5]])],
[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),
('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),
('history records the stage after advancing',
[[1, 5, 25, 100], [[0, 500], [0, 500]]],
[25, 'ramping', [5, 25]]),
('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),
('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),
('rollback after completion',
[[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],
[0, 'rolled_back', [50, 100, 0]]),
('guardrail sample 21',
[[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],
[100, 'complete', [5, 100, 100, 100, 100, 100]]),
('guardrail sample 23',
[[10, 50, 100], [[0.01, 500], [0.024, 99], [0, 50]]],
[50, 'ramping', [50, 50, 50]])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| low-sample breach only holds | [5, 'ramping', [1, 5]] | [5, 'ramping', [1, 5]] | Passed |
| rollback is terminal | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
| error rate exactly at threshold advances | [50, 'ramping', [50]] | [50, 'ramping', [50]] | Passed |
| error rate just above threshold rolls back | [0, 'rolled_back', [0]] | [0, 'rolled_back', [0]] | Passed |
| history records the stage after advancing | [25, 'ramping', [5, 25]] | [25, 'ramping', [5, 25]] | Passed |
| rollback after completion | [0, 'rolled_back', [50, 100, 0]] | [0, 'rolled_back', [50, 100, 0]] | Passed |
| guardrail sample 1 | [1, 'ramping', [1, 1, 1]] | [1, 'ramping', [1, 1, 1]] | Passed |
| guardrail sample 2 | [0, 'rolled_back', [0, 0, 0]] | [0, 'rolled_back', [0, 0, 0]] | Passed |
SHA-256 / a6a5aaa4d5c2ae39e050096b4708b95889b3250fa39a9daa0dde32ff7e0dcd9f
Verification & scope
A deterministic toy flag-evaluation model with a stipulated contract; it does not reproduce any vendor SDK byte for byte. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:54.561873+00:00.
Case digest / dd70059e25ff4580e448ad4de4dc730dc32bf39f31319d37f31dd9d38df80696