FAILURE MAP
← Case archive

FA-74396 / Experiment statistics / Open access

Welch unequal-variance t statistic: Welch df divides by arm sizes · case 01

Degrees of freedom are inflated, making small-sample p-values too small.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The df denominator uses (s^2/n)^2 / n instead of / (n - 1).

VERIFIED REPAIR

Use (va/na)^2/(na - 1) + (vb/nb)^2/(nb - 1).

Unsuccessful approach: Pooling the squared terms over na + nb - 2 is not the Satterthwaite form.

Case contract

a is control, b treatment. With sample variances (divisor n - 1), t = (mean_b - mean_a) / sqrt(va/na + vb/nb) and Welch-Satterthwaite df = (va/na + vb/nb)^2 / ((va/na)^2/(na-1) + (vb/nb)^2/(nb-1)). Fewer than two values in an arm or zero total variance -> None. Return [round(t, 6), round(df, 6)].

Why this case matters

Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na < 2 or nb < 2:
        return None
    ma, mb = sum(a) / na, sum(b) / nb
    va = sum((x - ma) ** 2 for x in a) / (na - 1)
    vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
    sa, sb = va / na, vb / nb
    if sa + sb == 0:
        return None
    t = (mb - ma) / math.sqrt(sa + sb)
    df = (sa + sb) ** 2 / (sa ** 2 / na + sb ** 2 / nb)
    return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
  ('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
 [('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
  ('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
  ('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
  ('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
  ('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 16',
   [[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
   [2.66816, 13.725637]),
  ('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
  ('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
  ('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
  ('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal sizes and variances[2.713602, 7.953743][2.713602, 6.594671]Failed
two-point arms[2.683282, 2.941176][2.683282, 1.470588]Failed
one arm constant[0.866025, 3.0][0.866025, 2.0]Failed
equal arms give zero t[0.0, 6.0][0.0, 4.0]Failed
negative effect[-7.778175, 5.468354][-7.778175, 4.075472]Failed
metric sample 1[1.290726, 9.302438][1.290726, 7.219744]Failed
metric sample 2[1.302114, 8.508486][1.302114, 6.806789]Failed
metric sample 3[2.622967, 4.443729][2.622967, 3.321892]Failed

SHA-256 / 3d5b4d10bad5ca65ae0e6e816ba775e33678f6ad2f43f669f47d97de47351c9b

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na < 2 or nb < 2:
        return None
    ma, mb = sum(a) / na, sum(b) / nb
    va = sum((x - ma) ** 2 for x in a) / (na - 1)
    vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
    sa, sb = va / na, vb / nb
    if sa + sb == 0:
        return None
    t = (mb - ma) / math.sqrt(sa + sb)
    df = (sa + sb) ** 2 / ((sa ** 2 + sb ** 2) / (na + nb - 2))
    return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
  ('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
 [('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
  ('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
  ('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
  ('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
  ('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 16',
   [[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
   [2.66816, 13.725637]),
  ('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
  ('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
  ('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
  ('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal sizes and variances[2.713602, 10.76885][2.713602, 6.594671]Failed
two-point arms[2.683282, 2.941176][2.683282, 1.470588]Failed
one arm constant[0.866025, 4.0][0.866025, 2.0]Failed
equal arms give zero t[0.0, 8.0][0.0, 4.0]Failed
negative effect[-7.778175, 6.923077][-7.778175, 4.075472]Failed
metric sample 1[1.290726, 15.925698][1.290726, 7.219744]Failed
metric sample 2[1.302114, 13.613578][1.302114, 6.806789]Failed
metric sample 3[2.622967, 4.458366][2.622967, 3.321892]Failed

SHA-256 / 499489c06bc45ef98be8d625c57d2a3ec22b47fe13c8d1e2cdfb639cf0c73a94

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na < 2 or nb < 2:
        return None
    ma, mb = sum(a) / na, sum(b) / nb
    va = sum((x - ma) ** 2 for x in a) / (na - 1)
    vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
    sa, sb = va / na, vb / nb
    if sa + sb == 0:
        return None
    t = (mb - ma) / math.sqrt(sa + sb)
    df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
    return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
  ('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
 [('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
  ('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
  ('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
  ('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
  ('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 16',
   [[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
   [2.66816, 13.725637]),
  ('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
  ('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
 [('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
  ('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
  ('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
  ('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
  ('single observation arm', [[1], [2, 3]], None),
  ('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
  ('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
  ('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal sizes and variances[2.713602, 6.594671][2.713602, 6.594671]Passed
two-point arms[2.683282, 1.470588][2.683282, 1.470588]Passed
one arm constant[0.866025, 2.0][0.866025, 2.0]Passed
equal arms give zero t[0.0, 4.0][0.0, 4.0]Passed
negative effect[-7.778175, 4.075472][-7.778175, 4.075472]Passed
metric sample 1[1.290726, 7.219744][1.290726, 7.219744]Passed
metric sample 2[1.302114, 6.806789][1.302114, 6.806789]Passed
metric sample 3[2.622967, 3.321892][2.622967, 3.321892]Passed

SHA-256 / 54bef7dfe4d4551012597af385bda3f1c469a2e21e99d12426b213657335d6a6

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.520725+00:00.

Case digest / 82f239fd2d88b6e726ac57540231e94c2557fb1349abd0b61af16d016021e227