{"abstract":"Users leaking into a disabled arm never trigger the mismatch alarm.","category":"Experiment statistics","checks":8,"contract":"counts and weights are per-arm lists. Expected count = total * w / sum(weights). chi2 sums (observed - expected)^2 / expected over positive-weight arms; any unit in a zero-weight arm is an immediate mismatch [None, True]. df = number of positive-weight arms - 1; mismatch iff chi2 exceeds the alpha = 0.001 critical value (10.828, 13.816, 16.266, 18.467, 20.515 for df 1..5). No units or df < 1 -> [0.0, False]. Return [round(chi2, 6), mismatch].","contract_signature":"counts, weights","evaluation_group":"w2-experiment-statistics-srm-chi-square","failed_approach":"Tolerating leakage up to half the traffic still hides small but real assignment bugs.","family":"w2-experiment-statistics-srm-chi-square-zero-weight-arm","id":"FA-74356","implementations":{"attempt":{"sha256":"876d16ebb27b8b47c887570ce54944d0bc4a7f76b6ef79e7e5cc728698ac8a30","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(counts, weights):\n    CRIT = {1: 10.828, 2: 13.816, 3: 16.266, 4: 18.467, 5: 20.515}\n    total = sum(counts)\n    wsum = sum(weights)\n    if total == 0:\n        return [0.0, False]\n    chi2 = 0.0\n    for o, w in zip(counts, weights):\n        if w == 0:\n            if o > total // 2:\n                return [None, True]\n            continue\n        e = total * w / wsum\n        chi2 += (o - e) ** 2 / e\n    df = sum(1 for w in weights if w > 0) - 1\n    if df < 1:\n        return [0.0, False]\n    return [round(chi2, 6), chi2 > CRIT[df]]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),\n  ('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),\n  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),\n  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),\n  ('arm count sample 2', [[45, 53, 73, 18], [1, 50, 2, 50]], [2400.801481, True])],\n [('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),\n  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),\n  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True]),\n  ('arm count sample 7', [[221, 4418, 96, 187], [2, 50, 1, 3]], [34.799989, True]),\n  ('arm count sample 47', [[4973, 62], [1, 0]], [None, True])],\n [('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),\n  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),\n  ('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),\n  ('arm count sample 12', [[45, 0, 60], [3, 1, 3]], [20.0, True]),\n  ('arm count sample 56', [[956, 26], [1, 0]], [None, True])],\n [('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),\n  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),\n  ('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),\n  ('arm count sample 17', [[2, 967], [1, 50]], [15.514737, True]),\n  ('arm count sample 18', [[95, 997, 0, 9], [1, 50, 1, 1]], [294.338365, True])],\n [('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),\n  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),\n  ('no units yet', [[0, 0], [1, 1]], [0.0, False]),\n  ('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),\n  ('arm count sample 22', [[18, 19, 157], [3, 2, 50]], [27.553608, True]),\n  ('arm count sample 47', [[4973, 62], [1, 0]], [None, True])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"broken":{"sha256":"718488e74f0474973d7c088b303f137d9d7586c71b9a974e2140f2b729e2e3ad","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\n\nN = 1\nobservations = []\ndef solve(counts, weights):\n    CRIT = {1: 10.828, 2: 13.816, 3: 16.266, 4: 18.467, 5: 20.515}\n    total = sum(counts)\n    wsum = sum(weights)\n    if total == 0:\n        return [0.0, False]\n    chi2 = 0.0\n    for o, w in zip(counts, weights):\n        if w == 0:\n            continue\n        e = total * w / wsum\n        chi2 += (o - e) ** 2 / e\n    df = sum(1 for w in weights if w > 0) - 1\n    if df < 1:\n        return [0.0, False]\n    return [round(chi2, 6), chi2 > CRIT[df]]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),\n  ('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),\n  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),\n  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),\n  ('arm count sample 2', [[45, 53, 73, 18], [1, 50, 2, 50]], [2400.801481, True])],\n [('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),\n  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),\n  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True]),\n  ('arm count sample 7', [[221, 4418, 96, 187], [2, 50, 1, 3]], [34.799989, True]),\n  ('arm count sample 47', [[4973, 62], [1, 0]], [None, True])],\n [('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),\n  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),\n  ('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),\n  ('arm count sample 12', [[45, 0, 60], [3, 1, 3]], [20.0, True]),\n  ('arm count sample 56', [[956, 26], [1, 0]], [None, True])],\n [('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),\n  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),\n  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),\n  ('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),\n  ('arm count sample 17', [[2, 967], [1, 50]], [15.514737, True]),\n  ('arm count sample 18', [[95, 997, 0, 9], [1, 50, 1, 1]], [294.338365, True])],\n [('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),\n  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),\n  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),\n  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),\n  ('no units yet', [[0, 0], [1, 1]], [0.0, False]),\n  ('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),\n  ('arm count sample 22', [[18, 19, 157], [3, 2, 50]], [27.553608, True]),\n  ('arm count sample 47', [[4973, 62], [1, 0]], [None, True])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"}},"limitations":"A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.","method":"Deterministic executable model with adversarial boundary fixtures.","provenance":{"created_by":"Failure Map","dependencies":"Python standard library","family":"w2-experiment-statistics-srm-chi-square-zero-weight-arm","generated_at":"2026-09-29T14:48:55.934978+00:00","license":"CC0-1.0","python":"3.12.14","seed":1,"split":"open-access"},"relevance":"SRM checks are the first gate on any experiment readout; a broken check hides assignment bugs.","root_cause":"Zero-weight arms are skipped without checking whether they received units.","sha256":"b92e4b836708586cc60254abe634163fd77ab0e7d2af9f6e38f835bb63000c40","title":"Sample ratio mismatch check: Units in a zero-weight arm are ignored · case 01","variant":1,"variant_policy":"Five numbered records share a model and may reuse boundary fixtures.","verified":true,"visibility":"public","verification":{"attempt":{"elapsed_ms":57.28,"exit_code":1,"observations":[{"actual":[0.64,false],"check":"balanced split within noise","expected":[0.64,false],"passed":true},{"actual":[0.0,false],"check":"unequal design weights are respected","expected":[0.0,false],"passed":true},{"actual":[0.4,false],"check":"ratio weights not in percent","expected":[0.4,false],"passed":true},{"actual":[16.0,true],"check":"clear mismatch at alpha 0.001","expected":[16.0,true],"passed":true},{"actual":[4.0,false],"check":"moderate imbalance below the strict threshold","expected":[4.0,false],"passed":true},{"actual":[0.008973,false],"check":"traffic in a zero-weight arm is a mismatch","expected":[null,true],"passed":false},{"actual":[0.263918,false],"check":"arm count sample 1","expected":[0.263918,false],"passed":true},{"actual":[2400.801481,true],"check":"arm count sample 2","expected":[2400.801481,true],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"balanced split within noise\", \"actual\": [0.64, false], \"expected\": [0.64, false], \"passed\": true}, {\"check\": \"unequal design weights are respected\", \"actual\": [0.0, false], \"expected\": [0.0, false], \"passed\": true}, {\"check\": \"ratio weights not in percent\", \"actual\": [0.4, false], \"expected\": [0.4, false], \"passed\": true}, {\"check\": \"clear mismatch at alpha 0.001\", \"actual\": [16.0, true], \"expected\": [16.0, true], \"passed\": true}, {\"check\": \"moderate imbalance below the strict threshold\", \"actual\": [4.0, false], \"expected\": [4.0, false], \"passed\": true}, {\"check\": \"traffic in a zero-weight arm is a mismatch\", \"actual\": [0.008973, false], \"expected\": [null, true], \"passed\": false}, {\"check\": \"arm count sample 1\", \"actual\": [0.263918, false], \"expected\": [0.263918, false], \"passed\": true}, {\"check\": \"arm count sample 2\", \"actual\": [2400.801481, true], \"expected\": [2400.801481, true], \"passed\": true}], \"passed\": false}\n"},"broken":{"elapsed_ms":43.23,"exit_code":1,"observations":[{"actual":[0.64,false],"check":"balanced split within noise","expected":[0.64,false],"passed":true},{"actual":[0.0,false],"check":"unequal design weights are respected","expected":[0.0,false],"passed":true},{"actual":[0.4,false],"check":"ratio weights not in percent","expected":[0.4,false],"passed":true},{"actual":[16.0,true],"check":"clear mismatch at alpha 0.001","expected":[16.0,true],"passed":true},{"actual":[4.0,false],"check":"moderate imbalance below the strict threshold","expected":[4.0,false],"passed":true},{"actual":[0.008973,false],"check":"traffic in a zero-weight arm is a mismatch","expected":[null,true],"passed":false},{"actual":[0.263918,false],"check":"arm count sample 1","expected":[0.263918,false],"passed":true},{"actual":[2400.801481,true],"check":"arm count sample 2","expected":[2400.801481,true],"passed":true}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"balanced split within noise\", \"actual\": [0.64, false], \"expected\": [0.64, false], \"passed\": true}, {\"check\": \"unequal design weights are respected\", \"actual\": [0.0, false], \"expected\": [0.0, false], \"passed\": true}, {\"check\": \"ratio weights not in percent\", \"actual\": [0.4, false], \"expected\": [0.4, false], \"passed\": true}, {\"check\": \"clear mismatch at alpha 0.001\", \"actual\": [16.0, true], \"expected\": [16.0, true], \"passed\": true}, {\"check\": \"moderate imbalance below the strict threshold\", \"actual\": [4.0, false], \"expected\": [4.0, false], \"passed\": true}, {\"check\": \"traffic in a zero-weight arm is a mismatch\", \"actual\": [0.008973, false], \"expected\": [null, true], \"passed\": false}, {\"check\": \"arm count sample 1\", \"actual\": [0.263918, false], \"expected\": [0.263918, false], \"passed\": true}, {\"check\": \"arm count sample 2\", \"actual\": [2400.801481, true], \"expected\": [2400.801481, true], \"passed\": true}], \"passed\": false}\n"}},"member_only":{"stages":["fixed"],"fields":["implementations.fixed","verification.fixed","harness","repair"],"note":"The verified repair, its recorded checks, the repair description, and the scoring harness are available to members."}}