FAILURE MAP
← Case archive

FA-74406 / Experiment statistics / Open access

Welch unequal-variance t statistic: Welch df numerator uses raw variances · case 01

Degrees of freedom scale with the metric's magnitude and sample size in nonsensical ways.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The numerator squares va + vb rather than va/na + vb/nb.

VERIFIED REPAIR

Square the sum of the per-arm variances of the mean.

Unsuccessful approach: Summing the squares instead of squaring the sum understates df.

Case contract

a is control, b treatment. With sample variances (divisor n - 1), t = (mean_b - mean_a) / sqrt(va/na + vb/nb) and Welch-Satterthwaite df = (va/na + vb/nb)^2 / ((va/na)^2/(na-1) + (vb/nb)^2/(nb-1)). Fewer than two values in an arm or zero total variance -> None. Return [round(t, 6), round(df, 6)].

Why this case matters

Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na < 2 or nb < 2:
        return None
    ma, mb = sum(a) / na, sum(b) / nb
    va = sum((x - ma) ** 2 for x in a) / (na - 1)
    vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
    sa, sb = va / na, vb / nb
    if sa + sb == 0:
        return None
    t = (mb - ma) / math.sqrt(sa + sb)
    df = (va + vb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
    return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
  ('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
 [('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
  ('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
  ('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
  ('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
  ('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 16',
   [[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
   [2.66816, 13.725637]),
  ('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
  ('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
  ('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354]),
  ('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal sizes and variances[2.713602, 214.033105][2.713602, 6.594671]Failed
two-point arms[2.683282, 5.882353][2.683282, 1.470588]Failed
one arm constant[0.866025, 18.0][0.866025, 2.0]Failed
equal arms give zero t[0.0, 36.0][0.0, 4.0]Failed
negative effect[-7.778175, 59.886792][-7.778175, 4.075472]Failed
metric sample 1[1.290726, 175.595848][1.290726, 7.219744]Failed
metric sample 2[1.302114, 170.169719][1.302114, 6.806789]Failed
metric sample 3[2.622967, 50.30028][2.622967, 3.321892]Failed

SHA-256 / e6cdd66e37cff6a4678efa70970de5e9ec89e4682fe578c0b0dbdc9bc62329b3

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na < 2 or nb < 2:
        return None
    ma, mb = sum(a) / na, sum(b) / nb
    va = sum((x - ma) ** 2 for x in a) / (na - 1)
    vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
    sa, sb = va / na, vb / nb
    if sa + sb == 0:
        return None
    t = (mb - ma) / math.sqrt(sa + sb)
    df = (sa ** 2 + sb ** 2) / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
    return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
  ('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
 [('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
  ('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
  ('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
  ('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
  ('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 16',
   [[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
   [2.66816, 13.725637]),
  ('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
  ('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
  ('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354]),
  ('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal sizes and variances[2.713602, 4.899071][2.713602, 6.594671]Failed
two-point arms[2.683282, 1.0][2.683282, 1.470588]Failed
one arm constant[0.866025, 2.0][0.866025, 2.0]Passed
equal arms give zero t[0.0, 2.0][0.0, 4.0]Failed
negative effect[-7.778175, 2.943396][-7.778175, 4.075472]Failed
metric sample 1[1.290726, 3.626714][1.290726, 7.219744]Failed
metric sample 2[1.302114, 4.0][1.302114, 6.806789]Failed
metric sample 3[2.622967, 2.980367][2.622967, 3.321892]Failed

SHA-256 / 15422d7b929696a776ab5890bc5977a312be0ab73d3349718a58c93104666a28

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na < 2 or nb < 2:
        return None
    ma, mb = sum(a) / na, sum(b) / nb
    va = sum((x - ma) ** 2 for x in a) / (na - 1)
    vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
    sa, sb = va / na, vb / nb
    if sa + sb == 0:
        return None
    t = (mb - ma) / math.sqrt(sa + sb)
    df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
    return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
  ('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
 [('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
  ('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
  ('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
  ('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
  ('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 16',
   [[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
   [2.66816, 13.725637]),
  ('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
  ('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
  ('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354]),
  ('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal sizes and variances[2.713602, 6.594671][2.713602, 6.594671]Passed
two-point arms[2.683282, 1.470588][2.683282, 1.470588]Passed
one arm constant[0.866025, 2.0][0.866025, 2.0]Passed
equal arms give zero t[0.0, 4.0][0.0, 4.0]Passed
negative effect[-7.778175, 4.075472][-7.778175, 4.075472]Passed
metric sample 1[1.290726, 7.219744][1.290726, 7.219744]Passed
metric sample 2[1.302114, 6.806789][1.302114, 6.806789]Passed
metric sample 3[2.622967, 3.321892][2.622967, 3.321892]Passed

SHA-256 / 0047f8c5b478366121b9a70d496b67c2a2f83aac6329bde0f3d36872b87319e6

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.574777+00:00.

Case digest / 28b86439cdfc9aedb2411d421bbafa43571ce9bd42f2f63cfeffb3afbd378d52