{"abstract":"Audit history shows the percentage before each advance, hiding when the ramp reached full traffic.","category":"Feature flag rollout bucketing","checks":8,"contract":"stages is an increasing list of percentages; the ramp starts at stages[0]. Each check [error_rate, sample] is processed in order: after a rollback everything stays at 0; a sample below 100 holds the current stage; an error rate above 0.02 rolls back to 0 permanently; otherwise advance one stage (staying at the last). State is complete at the last stage, rolled_back after a rollback, else ramping. Return [final percent, state, percent after each check].","evaluation_group":"w2-feature-flag-rollout-bucketing-guarded-ramp","failed_approach":"Appending the next stage (one ahead of the advance) double counts progress.","family":"w2-feature-flag-rollout-bucketing-guarded-ramp-history-timing","id":"FA-74211","implementations":{"attempt":{"sha256":"b75c2cd3ad194ab25fd6361b37595d2a86f50eeabfd1b3c281ab87a04aadec1f","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if state == 'rolled_back':\n            history.append(0)\n            continue\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if rate > 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n        history.append(stages[min(idx + 1, len(stages) - 1)])\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 20', [[1, 5, 25, 100], [[0, 100], [0, 100], [0, 50]]], [25, 'ramping', [5, 25, 25]]),\n  ('guardrail sample 24',\n   [[1, 5, 25, 100], [[0.01, 50], [0.01, 100], [0.02, 500], [0.024, 99], [0.01, 99], [0, 50]]],\n   [25, 'ramping', [1, 5, 25, 25, 25, 25]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 32',\n   [[10, 50, 100], [[0.01, 99], [0.01, 500], [0, 500], [0.02, 100], [0.024, 50], [0.02, 500]]],\n   [100, 'complete', [10, 50, 100, 100, 100, 100]]),\n  ('guardrail sample 39',\n   [[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],\n   [0, 'rolled_back', [50, 0, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),\n  ('guardrail sample 59',\n   [[1, 5, 25, 100], [[0.02, 50], [0.01, 100], [0.024, 50], [0.024, 50], [0.05, 99], [0.05, 50]]],\n   [5, 'ramping', [1, 5, 5, 5, 5, 5]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 23',\n   [[10, 50, 100], [[0.01, 500], [0.024, 99], [0, 50]]],\n   [50, 'ramping', [50, 50, 50]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"broken":{"sha256":"2663a132b8d0eb6918f160a131f0a3905a5746434b09ab5a714e24e413bce454","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if state == 'rolled_back':\n            history.append(0)\n            continue\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if rate > 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        history.append(stages[idx])\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 20', [[1, 5, 25, 100], [[0, 100], [0, 100], [0, 50]]], [25, 'ramping', [5, 25, 25]]),\n  ('guardrail sample 24',\n   [[1, 5, 25, 100], [[0.01, 50], [0.01, 100], [0.02, 500], [0.024, 99], [0.01, 99], [0, 50]]],\n   [25, 'ramping', [1, 5, 25, 25, 25, 25]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 32',\n   [[10, 50, 100], [[0.01, 99], [0.01, 500], [0, 500], [0.02, 100], [0.024, 50], [0.02, 500]]],\n   [100, 'complete', [10, 50, 100, 100, 100, 100]]),\n  ('guardrail sample 39',\n   [[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],\n   [0, 'rolled_back', [50, 0, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),\n  ('guardrail sample 59',\n   [[1, 5, 25, 100], [[0.02, 50], [0.01, 100], [0.024, 50], [0.024, 50], [0.05, 99], [0.05, 50]]],\n   [5, 'ramping', [1, 5, 5, 5, 5, 5]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 23',\n   [[10, 50, 100], [[0.01, 500], [0.024, 99], [0, 50]]],\n   [50, 'ramping', [50, 50, 50]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"fixed":{"sha256":"a6a5aaa4d5c2ae39e050096b4708b95889b3250fa39a9daa0dde32ff7e0dcd9f","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(stages, checks):\n    idx = 0\n    history = []\n    state = 'ramping' if len(stages) > 1 else 'complete'\n    for rate, sample in checks:\n        if state == 'rolled_back':\n            history.append(0)\n            continue\n        if sample < 100:\n            history.append(stages[idx])\n            continue\n        if rate > 0.02:\n            state = 'rolled_back'\n            history.append(0)\n            continue\n        idx = min(idx + 1, len(stages) - 1)\n        if idx == len(stages) - 1:\n            state = 'complete'\n        history.append(stages[idx])\n    final = 0 if state == 'rolled_back' else stages[idx]\n    return [final, state, history]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 1', [[1, 5, 25, 100], [[0.01, 50], [0.02, 50], [0.024, 50]]], [1, 'ramping', [1, 1, 1]]),\n  ('guardrail sample 2',\n   [[1, 5, 25, 100], [[0.024, 100], [0.01, 99], [0.05, 500]]],\n   [0, 'rolled_back', [0, 0, 0]])],\n [('rollback is terminal',\n   [[1, 5, 25, 100], [[0.05, 500], [0, 500], [0, 500]]],\n   [0, 'rolled_back', [0, 0, 0]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 20', [[1, 5, 25, 100], [[0, 100], [0, 100], [0, 50]]], [25, 'ramping', [5, 25, 25]]),\n  ('guardrail sample 24',\n   [[1, 5, 25, 100], [[0.01, 50], [0.01, 100], [0.02, 500], [0.024, 99], [0.01, 99], [0, 50]]],\n   [25, 'ramping', [1, 5, 25, 25, 25, 25]])],\n [('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('guardrail sample 11',\n   [[100], [[0.05, 100], [0.024, 50], [0.02, 50], [0.02, 99], [0.05, 50], [0.05, 50]]],\n   [0, 'rolled_back', [0, 0, 0, 0, 0, 0]]),\n  ('guardrail sample 32',\n   [[10, 50, 100], [[0.01, 99], [0.01, 500], [0, 500], [0.02, 100], [0.024, 50], [0.02, 500]]],\n   [100, 'complete', [10, 50, 100, 100, 100, 100]]),\n  ('guardrail sample 39',\n   [[10, 50, 100], [[0, 500], [0.024, 500], [0.02, 99], [0.024, 99]]],\n   [0, 'rolled_back', [50, 0, 0, 0]])],\n [('error rate just above threshold rolls back', [[10, 50, 100], [[0.024, 500]]], [0, 'rolled_back', [0]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 16',\n   [[5, 100], [[0.01, 50], [0.02, 50], [0.024, 100], [0.02, 50], [0.024, 500]]],\n   [0, 'rolled_back', [5, 5, 0, 0, 0]]),\n  ('guardrail sample 46', [[5, 100], [[0.01, 50], [0.02, 100]]], [100, 'complete', [5, 100]]),\n  ('guardrail sample 59',\n   [[1, 5, 25, 100], [[0.02, 50], [0.01, 100], [0.024, 50], [0.024, 50], [0.05, 99], [0.05, 50]]],\n   [5, 'ramping', [1, 5, 5, 5, 5, 5]])],\n [('low-sample breach only holds', [[1, 5, 25, 100], [[0.5, 20], [0, 500]]], [5, 'ramping', [1, 5]]),\n  ('error rate exactly at threshold advances', [[10, 50, 100], [[0.02, 500]]], [50, 'ramping', [50]]),\n  ('history records the stage after advancing',\n   [[1, 5, 25, 100], [[0, 500], [0, 500]]],\n   [25, 'ramping', [5, 25]]),\n  ('final percent after rollback is zero', [[5, 100], [[0, 500], [0.1, 500]]], [0, 'rolled_back', [100, 0]]),\n  ('single stage is complete from the start', [[100], [[0, 50]]], [100, 'complete', [100]]),\n  ('rollback after completion',\n   [[10, 50, 100], [[0, 500], [0, 500], [0.03, 500]]],\n   [0, 'rolled_back', [50, 100, 0]]),\n  ('guardrail sample 21',\n   [[5, 100], [[0.024, 99], [0, 100], [0.02, 99], [0.01, 50], [0.02, 50], [0.02, 100]]],\n   [100, 'complete', [5, 100, 100, 100, 100, 100]]),\n  ('guardrail sample 23',\n   [[10, 50, 100], [[0.01, 500], [0.024, 99], [0, 50]]],\n   [50, 'ramping', [50, 50, 50]])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"}},"limitations":"A deterministic toy flag-evaluation model with a stipulated contract; it does not reproduce any vendor SDK byte for byte. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.","method":"Deterministic executable model with adversarial boundary fixtures.","provenance":{"created_by":"Failure Map","dependencies":"Python standard library","family":"w2-feature-flag-rollout-bucketing-guarded-ramp-history-timing","generated_at":"2026-09-29T14:48:54.561873+00:00","license":"CC0-1.0","python":"3.12.14","seed":1,"split":"open-access"},"relevance":"Automated ramps with guardrails must neither overreact to noise nor resume after a rollback.","repair":"Append the stage after advancing.","root_cause":"The stage is appended to history before idx advances.","sha256":"dd70059e25ff4580e448ad4de4dc730dc32bf39f31319d37f31dd9d38df80696","title":"Guarded progressive ramp: History lags one stage behind · case 01","variant":1,"variant_policy":"Five numbered records share a model and may reuse boundary fixtures.","verification":{"attempt":{"elapsed_ms":38.851,"exit_code":1,"observations":[{"actual":[5,"ramping",[1,25]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":false},{"actual":[0,"rolled_back",[0,0,0]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[50,"ramping",[100]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":false},{"actual":[0,"rolled_back",[0]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":true},{"actual":[25,"ramping",[25,100]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":false},{"actual":[0,"rolled_back",[100,100,0]],"check":"rollback after completion","expected":[0,"rolled_back",[50,100,0]],"passed":false},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 25]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": false}, {\"check\": \"rollback is terminal\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [50, \"ramping\", [100]], \"expected\": [50, \"ramping\", [50]], \"passed\": false}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": true}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [25, 100]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": false}, {\"check\": \"rollback after completion\", \"actual\": [0, \"rolled_back\", [100, 100, 0]], \"expected\": [0, \"rolled_back\", [50, 100, 0]], \"passed\": false}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}], \"passed\": false}\n"},"broken":{"elapsed_ms":40.717,"exit_code":1,"observations":[{"actual":[5,"ramping",[1,1]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":false},{"actual":[0,"rolled_back",[0,0,0]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[50,"ramping",[10]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":false},{"actual":[0,"rolled_back",[0]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":true},{"actual":[25,"ramping",[1,5]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":false},{"actual":[0,"rolled_back",[10,50,0]],"check":"rollback after completion","expected":[0,"rolled_back",[50,100,0]],"passed":false},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 1]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": false}, {\"check\": \"rollback is terminal\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [50, \"ramping\", [10]], \"expected\": [50, \"ramping\", [50]], \"passed\": false}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": true}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [1, 5]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": false}, {\"check\": \"rollback after completion\", \"actual\": [0, \"rolled_back\", [10, 50, 0]], \"expected\": [0, \"rolled_back\", [50, 100, 0]], \"passed\": false}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}], \"passed\": false}\n"},"fixed":{"elapsed_ms":41.445,"exit_code":0,"observations":[{"actual":[5,"ramping",[1,5]],"check":"low-sample breach only holds","expected":[5,"ramping",[1,5]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"rollback is terminal","expected":[0,"rolled_back",[0,0,0]],"passed":true},{"actual":[50,"ramping",[50]],"check":"error rate exactly at threshold advances","expected":[50,"ramping",[50]],"passed":true},{"actual":[0,"rolled_back",[0]],"check":"error rate just above threshold rolls back","expected":[0,"rolled_back",[0]],"passed":true},{"actual":[25,"ramping",[5,25]],"check":"history records the stage after advancing","expected":[25,"ramping",[5,25]],"passed":true},{"actual":[0,"rolled_back",[50,100,0]],"check":"rollback after completion","expected":[0,"rolled_back",[50,100,0]],"passed":true},{"actual":[1,"ramping",[1,1,1]],"check":"guardrail sample 1","expected":[1,"ramping",[1,1,1]],"passed":true},{"actual":[0,"rolled_back",[0,0,0]],"check":"guardrail sample 2","expected":[0,"rolled_back",[0,0,0]],"passed":true}],"passed":true,"stderr":"","stdout":"{\"observations\": [{\"check\": \"low-sample breach only holds\", \"actual\": [5, \"ramping\", [1, 5]], \"expected\": [5, \"ramping\", [1, 5]], \"passed\": true}, {\"check\": \"rollback is terminal\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}, {\"check\": \"error rate exactly at threshold advances\", \"actual\": [50, \"ramping\", [50]], \"expected\": [50, \"ramping\", [50]], \"passed\": true}, {\"check\": \"error rate just above threshold rolls back\", \"actual\": [0, \"rolled_back\", [0]], \"expected\": [0, \"rolled_back\", [0]], \"passed\": true}, {\"check\": \"history records the stage after advancing\", \"actual\": [25, \"ramping\", [5, 25]], \"expected\": [25, \"ramping\", [5, 25]], \"passed\": true}, {\"check\": \"rollback after completion\", \"actual\": [0, \"rolled_back\", [50, 100, 0]], \"expected\": [0, \"rolled_back\", [50, 100, 0]], \"passed\": true}, {\"check\": \"guardrail sample 1\", \"actual\": [1, \"ramping\", [1, 1, 1]], \"expected\": [1, \"ramping\", [1, 1, 1]], \"passed\": true}, {\"check\": \"guardrail sample 2\", \"actual\": [0, \"rolled_back\", [0, 0, 0]], \"expected\": [0, \"rolled_back\", [0, 0, 0]], \"passed\": true}], \"passed\": true}\n"}},"verified":true,"visibility":"public"}