{"abstract":"Launches sitting exactly at the tolerated 2 percent error rate are rolled back.","category":"Feature flag rollout bucketing","checks":8,"contract":"stages is an increasing list of percentages; the ramp starts at stages[0]. Each check [error_rate, sample] is processed in order: after a rollback everything stays at 0; a sample below 100 holds the current stage; an error rate above 0.02 rolls back to 0 permanently; otherwise advance one stage (staying at the last). State is complete at the last stage, rolled_back after a rollback, else ramping. Return [final percent, state, percent after each check].","contract_signature":"stages, checks","evaluation_group":"w2-feature-flag-rollout-bucketing-guarded-ramp","failed_approach":"Rounding the rate to two decimals first hides breaches such as 0.024.","family":"w2-feature-flag-rollout-bucketing-guarded-ramp-error-threshold","id":"FA-74206","implementations":{"attempt":{"sha256":"318f0796ef2933b4742ad56692df980fc7faca787937d54e1b383a74968db0bc","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if state == 'rolled_back':\n            history.append(0)\n            continue\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if round(rate, 2) > 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n        history.append(stages[idx])\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('guardrail sample 6',\n   [[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),\n  ('guardrail sample 28',\n   [[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0]]),\n  ('guardrail sample 31',\n   [[100], [[0.05, 99], [0.01, 50], [0.02, 500], [0.02, 100], [0.01, 500], [0.05, 500]]],\n   [0, 'rolled_back', [100, 100, 100, 100, 100, 0]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),\n  ('guardrail sample 51',\n   [[1, 5, 25, 100], [[0.05, 99], [0.05, 99], [0.02, 100], [0.05, 99], [0.024, 500], [0.01, 100]]],\n   [0, 'rolled_back', [1, 1, 5, 5, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 18',\n   [[10, 50, 100], [[0.024, 500], [0.05, 50], [0.02, 100], [0.02, 500], [0.05, 500], [0.024, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 28',\n   [[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 22',\n   [[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),\n  ('guardrail sample 40',\n   [[1, 5, 25, 100], [[0.02, 100], [0.01, 500], [0.024, 100], [0, 99], [0.02, 500], [0.01, 500]]],\n   [0, 'rolled_back', [5, 25, 0, 0, 0, 0]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"broken":{"sha256":"9982017b58f7ba6132892684c60844431724db3f303a3e4c963495b831342eef","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if state == 'rolled_back':\n            history.append(0)\n            continue\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if rate >= 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n        history.append(stages[idx])\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('guardrail sample 3', [[1, 5, 25, 100], [[0.05, 99], [0.05, 500]]], [0, 'rolled_back', [1, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('guardrail sample 6',\n   [[100], [[0.024, 99], [0.01, 99], [0.02, 500], [0.024, 500], [0.01, 50], [0.024, 100]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),\n  ('guardrail sample 28',\n   [[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0]]),\n  ('guardrail sample 31',\n   [[100], [[0.05, 99], [0.01, 50], [0.02, 500], [0.02, 100], [0.01, 500], [0.05, 500]]],\n   [0, 'rolled_back', [100, 100, 100, 100, 100, 0]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),\n  ('guardrail sample 51',\n   [[1, 5, 25, 100], [[0.05, 99], [0.05, 99], [0.02, 100], [0.05, 99], [0.024, 500], [0.01, 100]]],\n   [0, 'rolled_back', [1, 1, 5, 5, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 18',\n   [[10, 50, 100], [[0.024, 500], [0.05, 50], [0.02, 100], [0.02, 500], [0.05, 500], [0.024, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 28',\n   [[5, 100], [[0.02, 500], [0.01, 500], [0.01, 50], [0.024, 500], [0.024, 500]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 22',\n   [[100], [[0.01, 100], [0.05, 99], [0.01, 100], [0.05, 100], [0, 500], [0.01, 50]]],\n   [0, 'rolled_back', [100, 100, 100, 0, 0, 0]]),\n  ('guardrail sample 40',\n   [[1, 5, 25, 100], [[0.02, 100], [0.01, 500], [0.024, 100], [0, 99], [0.02, 500], [0.01, 500]]],\n   [0, 'rolled_back', [5, 25, 0, 0, 0, 0]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"}},"limitations":"A deterministic toy flag-evaluation model with a stipulated contract; it does not reproduce any vendor SDK byte for byte. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.","method":"Deterministic executable model with adversarial boundary fixtures.","provenance":{"created_by":"Failure Map","dependencies":"Python standard library","family":"w2-feature-flag-rollout-bucketing-guarded-ramp-error-threshold","generated_at":"2026-09-29T14:48:54.559669+00:00","license":"CC0-1.0","python":"3.12.14","seed":1,"split":"open-access"},"relevance":"Automated ramps with guardrails must neither overreact to noise nor resume after a rollback.","root_cause":"The breach test uses rate >= 0.02.","sha256":"987362eeee74460b7114df7279e805ea43cfd4a282b2d332f2713efcf87c93d1","title":"Guarded progressive ramp: An error rate equal to the threshold rolls back · case 01","variant":1,"variant_policy":"Five numbered records share a model and may reuse boundary fixtures.","verified":true,"visibility":"public","verification":{"attempt":{"elapsed_ms":39.941,"exit_code":1,"observations":[{"actual":[5,"ramping",[1,5]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[50,"ramping",[50]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":true},{"actual":[50,"ramping",[50]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":false},{"actual":[25,"ramping",[5,25]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":true},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[5,5,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":false},{"actual":[0,"rolled_back",[1,0]],"check":"guardrail sample 3","expected":[0,"rolled_back",[1,0]],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 5]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": true}, {\"check\": \"rollback is terminal\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [50, \"ramping\", [50]], \"expected\": [50, \"ramping\", [50]], \"passed\": true}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [50, \"ramping\", [50]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": false}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [5, 25]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": true}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [5, 5, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": false}, {\"check\": \"guardrail sample 3\", \"actual\": [0, \"rolled_back\", [1, 0]], \"expected\": [0, \"rolled_back\", [1, 0]], \"passed\": true}], \"passed\": false}\n"},"broken":{"elapsed_ms":41.415,"exit_code":1,"observations":[{"actual":[5,"ramping",[1,5]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[0,"rolled_back",[0]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":false},{"actual":[0,"rolled_back",[0]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":true},{"actual":[25,"ramping",[5,25]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":true},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[0,"rolled_back",[1,0]],"check":"guardrail sample 3","expected":[0,"rolled_back",[1,0]],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 5]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": true}, {\"check\": \"rollback is terminal\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [50, \"ramping\", [50]], \"passed\": false}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": true}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [5, 25]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": true}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"guardrail sample 3\", \"actual\": [0, \"rolled_back\", [1, 0]], \"expected\": [0, \"rolled_back\", [1, 0]], \"passed\": true}], \"passed\": false}\n"}},"member_only":{"stages":["fixed"],"fields":["implementations.fixed","verification.fixed","harness","repair"],"note":"The verified repair, its recorded checks, the repair description, and the scoring harness are available to members."}}