FAILURE MAP
← Case archive

FA-74346 / Experiment statistics / Open access

Sample ratio mismatch check: Design weights are assumed to be equal · case 01

A deliberate 2:1 allocation is flagged as a sample ratio mismatch on every readout.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

Expected counts are total / number of arms, ignoring the configured weights.

VERIFIED REPAIR

Compute expected counts from total * w / sum(weights).

Unsuccessful approach: Dividing by 100 assumes weights are percentages and breaks ratio-style weights.

Case contract

counts and weights are per-arm lists. Expected count = total * w / sum(weights). chi2 sums (observed - expected)^2 / expected over positive-weight arms; any unit in a zero-weight arm is an immediate mismatch [None, True]. df = number of positive-weight arms - 1; mismatch iff chi2 exceeds the alpha = 0.001 critical value (10.828, 13.816, 16.266, 18.467, 20.515 for df 1..5). No units or df < 1 -> [0.0, False]. Return [round(chi2, 6), mismatch].

Why this case matters

SRM checks are the first gate on any experiment readout; a broken check hides assignment bugs.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(counts, weights):
    CRIT = {1: 10.828, 2: 13.816, 3: 16.266, 4: 18.467, 5: 20.515}
    total = sum(counts)
    wsum = sum(weights)
    if total == 0:
        return [0.0, False]
    chi2 = 0.0
    for o, w in zip(counts, weights):
        if w == 0:
            if o > 0:
                return [None, True]
            continue
        e = total / len(counts)
        chi2 += (o - e) ** 2 / e
    df = sum(1 for w in weights if w > 0) - 1
    if df < 1:
        return [0.0, False]
    return [round(chi2, 6), chi2 > CRIT[df]]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),
  ('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('three arms uneven', [[3100, 3300, 3600], [1, 1, 1]], [38.0, True]),
  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False])],
 [('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('three arms uneven', [[3100, 3300, 3600], [1, 1, 1]], [38.0, True]),
  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),
  ('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True])],
 [('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('arm count sample 8', [[0, 0, 152, 0], [0, 1, 50, 1]], [6.08, False]),
  ('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),
  ('arm count sample 14', [[4851, 168], [50, 1]], [50.190879, True])],
 [('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('arm count sample 15', [[72, 0, 108], [3, 0, 50]], [397.488, True]),
  ('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),
  ('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])],
 [('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('no units yet', [[0, 0], [1, 1]], [0.0, False]),
  ('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),
  ('arm count sample 25', [[26, 4871], [1, 50]], [52.081033, True]),
  ('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
balanced split within noise[0.64, False][0.64, False]Passed
unequal design weights are respected[333.333333, True][0.0, False]Failed
ratio weights not in percent[0.4, False][0.4, False]Passed
clear mismatch at alpha 0.001[16.0, True][16.0, True]Passed
moderate imbalance below the strict threshold[4.0, False][4.0, False]Passed
empty zero-weight arm is ignored[169.066667, True][1.6, False]Failed
three arms uneven[38.0, True][38.0, True]Passed
arm count sample 1[0.263918, False][0.263918, False]Passed

SHA-256 / f4cd41baf7378040f7a8fca41e958c34b288f928a1928b99fd05ecf36b49f8d5

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(counts, weights):
    CRIT = {1: 10.828, 2: 13.816, 3: 16.266, 4: 18.467, 5: 20.515}
    total = sum(counts)
    wsum = sum(weights)
    if total == 0:
        return [0.0, False]
    chi2 = 0.0
    for o, w in zip(counts, weights):
        if w == 0:
            if o > 0:
                return [None, True]
            continue
        e = total * w / 100
        chi2 += (o - e) ** 2 / e
    df = sum(1 for w in weights if w > 0) - 1
    if df < 1:
        return [0.0, False]
    return [round(chi2, 6), chi2 > CRIT[df]]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),
  ('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('three arms uneven', [[3100, 3300, 3600], [1, 1, 1]], [38.0, True]),
  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False])],
 [('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('three arms uneven', [[3100, 3300, 3600], [1, 1, 1]], [38.0, True]),
  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),
  ('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True])],
 [('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('arm count sample 8', [[0, 0, 152, 0], [0, 1, 50, 1]], [6.08, False]),
  ('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),
  ('arm count sample 14', [[4851, 168], [50, 1]], [50.190879, True])],
 [('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('arm count sample 15', [[72, 0, 108], [3, 0, 50]], [397.488, True]),
  ('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),
  ('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])],
 [('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('no units yet', [[0, 0], [1, 1]], [0.0, False]),
  ('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),
  ('arm count sample 25', [[26, 4871], [1, 50]], [52.081033, True]),
  ('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
balanced split within noise[0.64, False][0.64, False]Passed
unequal design weights are respected[94090.0, True][0.0, False]Failed
ratio weights not in percent[48040.0, True][0.4, False]Failed
clear mismatch at alpha 0.001[481000.0, True][16.0, True]Failed
moderate imbalance below the strict threshold[480400.0, True][4.0, False]Failed
empty zero-weight arm is ignored[48100.0, True][1.6, False]Failed
three arms uneven[314900.0, True][38.0, True]Failed
arm count sample 1[14289.265292, True][0.263918, False]Failed

SHA-256 / 5fb98384b06b30048ebba78dcf16ae94d92ec217fa402cf6cbdbc6bc4aa50eb9

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(counts, weights):
    CRIT = {1: 10.828, 2: 13.816, 3: 16.266, 4: 18.467, 5: 20.515}
    total = sum(counts)
    wsum = sum(weights)
    if total == 0:
        return [0.0, False]
    chi2 = 0.0
    for o, w in zip(counts, weights):
        if w == 0:
            if o > 0:
                return [None, True]
            continue
        e = total * w / wsum
        chi2 += (o - e) ** 2 / e
    df = sum(1 for w in weights if w > 0) - 1
    if df < 1:
        return [0.0, False]
    return [round(chi2, 6), chi2 > CRIT[df]]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('balanced split within noise', [[5040, 4960], [50, 50]], [0.64, False]),
  ('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('three arms uneven', [[3100, 3300, 3600], [1, 1, 1]], [38.0, True]),
  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False])],
 [('unequal design weights are respected', [[2000, 1000], [2, 1]], [0.0, False]),
  ('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('three arms uneven', [[3100, 3300, 3600], [1, 1, 1]], [38.0, True]),
  ('arm count sample 1', [[493, 477], [3, 3]], [0.263918, False]),
  ('arm count sample 6', [[32, 944, 99], [1, 50, 2]], [95.793172, True])],
 [('ratio weights not in percent', [[510, 490], [1, 1]], [0.4, False]),
  ('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('arm count sample 8', [[0, 0, 152, 0], [0, 1, 50, 1]], [6.08, False]),
  ('arm count sample 11', [[92, 0], [2, 1]], [46.0, True]),
  ('arm count sample 14', [[4851, 168], [50, 1]], [50.190879, True])],
 [('clear mismatch at alpha 0.001', [[5200, 4800], [1, 1]], [16.0, True]),
  ('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('arm count sample 15', [[72, 0, 108], [3, 0, 50]], [397.488, True]),
  ('arm count sample 16', [[0, 0, 44, 31], [2, 0, 3, 50]], [412.339111, True]),
  ('arm count sample 23', [[99, 96], [2, 3]], [9.423077, False])],
 [('moderate imbalance below the strict threshold', [[5100, 4900], [1, 1]], [4.0, False]),
  ('traffic in a zero-weight arm is a mismatch', [[500, 500, 3], [1, 1, 0]], [None, True]),
  ('drained arm does not loosen the threshold', [[5170, 4830, 0], [1, 1, 0]], [11.56, True]),
  ('empty zero-weight arm is ignored', [[520, 480, 0], [1, 1, 0]], [1.6, False]),
  ('no units yet', [[0, 0], [1, 1]], [0.0, False]),
  ('arm count sample 21', [[67, 446, 551], [1, 50, 50]], [316.144117, True]),
  ('arm count sample 25', [[26, 4871], [1, 50]], [52.081033, True]),
  ('arm count sample 32', [[490, 265, 153, 180], [3, 2, 1, 1]], [11.893229, False])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
balanced split within noise[0.64, False][0.64, False]Passed
unequal design weights are respected[0.0, False][0.0, False]Passed
ratio weights not in percent[0.4, False][0.4, False]Passed
clear mismatch at alpha 0.001[16.0, True][16.0, True]Passed
moderate imbalance below the strict threshold[4.0, False][4.0, False]Passed
empty zero-weight arm is ignored[1.6, False][1.6, False]Passed
three arms uneven[38.0, True][38.0, True]Passed
arm count sample 1[0.263918, False][0.263918, False]Passed

SHA-256 / b82b4c515751496414a8d836693758d33016397f2a9e3bc729b1ac134be2585b

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:55.912283+00:00.

Case digest / ff14b7286a75630721935b2d349eb6b4fb2ed7eb55242c8e4e74bd57cb10054c