{"abstract":"After an automatic rollback the ramp quietly climbs back up without human review.","category":"Feature flag rollout bucketing","checks":8,"contract":"stages is an increasing list of percentages; the ramp starts at stages[0]. Each check [error_rate, sample] is processed in order: after a rollback everything stays at 0; a sample below 100 holds the current stage; an error rate above 0.02 rolls back to 0 permanently; otherwise advance one stage (staying at the last). State is complete at the last stage, rolled_back after a rollback, else ramping. Return [final percent, state, percent after each check].","evaluation_group":"w2-feature-flag-rollout-bucketing-guarded-ramp","failed_approach":"Resetting to the first stage on the next check still resumes the ramp automatically.","family":"w2-feature-flag-rollout-bucketing-guarded-ramp-rollback-terminal","id":"FA-74201","implementations":{"attempt":{"sha256":"be649478cd5f98bf41ae013c2af1b93ed8050fdba66fb4cce8dc0a8051d3edcd","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if state == 'rolled_back':\n            state = 'ramping'\n            idx = 0\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if rate > 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n        history.append(stages[idx])\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('guardrail sample 6',\n   [[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),\n  ('guardrail sample 7',\n   [[10, 50, 100], [[0, 50], [0, 50], [0.05, 99], [0.02, 50]]],\n   [10, 'ramping', [10, 10, 10, 10]]),\n  ('guardrail sample 19',\n   [[5, 100], [[0, 100], [0.05, 50], [0.024, 500], [0.01, 100], [0.024, 50]]],\n   [0, 'rolled_back', [100, 100, 0, 0, 0]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 12', [[10, 50, 100], [[0.02, 100]]], [50, 'ramping', [50]]),\n  ('guardrail sample 39',\n   [[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],\n   [0, 'rolled_back', [50, 0, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 17', [[5, 100], [[0.05, 500], [0.01, 99]]], [0, 'rolled_back', [0, 0]]),\n  ('guardrail sample 57',\n   [[1, 5, 25, 100], [[0.024, 500], [0.024, 99], [0.02, 99]]],\n   [0, 'rolled_back', [0, 0, 0]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 17', [[5, 100], [[0.05, 500], [0.01, 99]]], [0, 'rolled_back', [0, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 22',\n   [[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"broken":{"sha256":"18ff97c3404f4778b12e887982b1985017760ec4e90b29902f3b8d8015d5b7c5","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if rate > 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n        history.append(stages[idx])\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('guardrail sample 6',\n   [[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),\n  ('guardrail sample 7',\n   [[10, 50, 100], [[0, 50], [0, 50], [0.05, 99], [0.02, 50]]],\n   [10, 'ramping', [10, 10, 10, 10]]),\n  ('guardrail sample 19',\n   [[5, 100], [[0, 100], [0.05, 50], [0.024, 500], [0.01, 100], [0.024, 50]]],\n   [0, 'rolled_back', [100, 100, 0, 0, 0]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 12', [[10, 50, 100], [[0.02, 100]]], [50, 'ramping', [50]]),\n  ('guardrail sample 39',\n   [[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],\n   [0, 'rolled_back', [50, 0, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 17', [[5, 100], [[0.05, 500], [0.01, 99]]], [0, 'rolled_back', [0, 0]]),\n  ('guardrail sample 57',\n   [[1, 5, 25, 100], [[0.024, 500], [0.024, 99], [0.02, 99]]],\n   [0, 'rolled_back', [0, 0, 0]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 17', [[5, 100], [[0.05, 500], [0.01, 99]]], [0, 'rolled_back', [0, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 22',\n   [[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"fixed":{"sha256":"2108669da4fec1b7f8e8147367dc36657fd50e63adcce0df997c825b35d348bc","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if state == 'rolled_back':\n            history.append(0)\n            continue\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if rate > 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n        history.append(stages[idx])\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('guardrail sample 6',\n   [[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),\n  ('guardrail sample 7',\n   [[10, 50, 100], [[0, 50], [0, 50], [0.05, 99], [0.02, 50]]],\n   [10, 'ramping', [10, 10, 10, 10]]),\n  ('guardrail sample 19',\n   [[5, 100], [[0, 100], [0.05, 50], [0.024, 500], [0.01, 100], [0.024, 50]]],\n   [0, 'rolled_back', [100, 100, 0, 0, 0]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 12', [[10, 50, 100], [[0.02, 100]]], [50, 'ramping', [50]]),\n  ('guardrail sample 39',\n   [[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],\n   [0, 'rolled_back', [50, 0, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 17', [[5, 100], [[0.05, 500], [0.01, 99]]], [0, 'rolled_back', [0, 0]]),\n  ('guardrail sample 57',\n   [[1, 5, 25, 100], [[0.024, 500], [0.024, 99], [0.02, 99]]],\n   [0, 'rolled_back', [0, 0, 0]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 17', [[5, 100], [[0.05, 500], [0.01, 99]]], [0, 'rolled_back', [0, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 22',\n   [[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"}},"limitations":"A deterministic toy flag-evaluation model with a stipulated contract; it does not reproduce any vendor SDK byte for byte. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.","method":"Deterministic executable model with adversarial boundary fixtures.","provenance":{"created_by":"Failure Map","dependencies":"Python standard library","family":"w2-feature-flag-rollout-bucketing-guarded-ramp-rollback-terminal","generated_at":"2026-09-29T14:48:54.518653+00:00","license":"CC0-1.0","python":"3.12.14","seed":1,"split":"open-access"},"relevance":"Automated ramps with guardrails must neither overreact to noise nor resume after a rollback.","repair":"Keep the ramp at 0 for every check after a rollback.","root_cause":"Nothing short-circuits later checks once the state is rolled_back.","sha256":"a498f0cff60f114fe871fdba3a12aa7a19b1174e583856b906c4abf8ede7b205","title":"Guarded progressive ramp: Healthy intervals resume a rolled-back ramp · case 01","variant":1,"variant_policy":"Five numbered records share a model and may reuse boundary fixtures.","verification":{"attempt":{"elapsed_ms":38.639,"exit_code":1,"observations":[{"actual":[5,"ramping",[1,5]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":true},{"actual":[25,"ramping",[0,5,25]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":false},{"actual":[50,"ramping",[50]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":true},{"actual":[0,"rolled_back",[0]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":true},{"actual":[25,"ramping",[5,25]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":true},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[0,1,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":false},{"actual":[0,"rolled_back",[1,0]],"check":"guardrail sample 3","expected":[0,"rolled_back",[1,0]],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 5]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": true}, {\"check\": \"rollback is terminal\", \"actual\": [25, \"ramping\", [0, 5, 25]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": false}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [50, \"ramping\", [50]], \"expected\": [50, \"ramping\", [50]], \"passed\": true}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": true}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [5, 25]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": true}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [0, 1, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": false}, {\"check\": \"guardrail sample 3\", \"actual\": [0, \"rolled_back\", [1, 0]], \"expected\": [0, \"rolled_back\", [1, 0]], \"passed\": true}], \"passed\": false}\n"},"broken":{"elapsed_ms":44.914,"exit_code":1,"observations":[{"actual":[5,"ramping",[1,5]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":true},{"actual":[0,"rolled_back",[0,5,25]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":false},{"actual":[50,"ramping",[50]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":true},{"actual":[0,"rolled_back",[0]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":true},{"actual":[25,"ramping",[5,25]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":true},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[0,1,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":false},{"actual":[0,"rolled_back",[1,0]],"check":"guardrail sample 3","expected":[0,"rolled_back",[1,0]],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 5]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": true}, {\"check\": \"rollback is terminal\", \"actual\": [0, \"rolled_back\", [0, 5, 25]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": false}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [50, \"ramping\", [50]], \"expected\": [50, \"ramping\", [50]], \"passed\": true}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": true}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [5, 25]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": true}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [0, 1, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": false}, {\"check\": \"guardrail sample 3\", \"actual\": [0, \"rolled_back\", [1, 0]], \"expected\": [0, \"rolled_back\", [1, 0]], \"passed\": true}], \"passed\": false}\n"},"fixed":{"elapsed_ms":39.314,"exit_code":0,"observations":[{"actual":[5,"ramping",[1,5]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[50,"ramping",[50]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":true},{"actual":[0,"rolled_back",[0]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":true},{"actual":[25,"ramping",[5,25]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":true},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[0,"rolled_back",[1,0]],"check":"guardrail sample 3","expected":[0,"rolled_back",[1,0]],"passed":true}],"passed":true,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 5]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": true}, {\"check\": \"rollback is terminal\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [50, \"ramping\", [50]], \"expected\": [50, \"ramping\", [50]], \"passed\": true}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": true}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [5, 25]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": true}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"guardrail sample 3\", \"actual\": [0, \"rolled_back\", [1, 0]], \"expected\": [0, \"rolled_back\", [1, 0]], \"passed\": true}], \"passed\": true}\n"}},"verified":true,"visibility":"public"}