FA-74396 / Experiment statistics / Open access
Welch unequal-variance t statistic: Welch df divides by arm sizes · case 01
Degrees of freedom are inflated, making small-sample p-values too small.
ROOT CAUSE
The df denominator uses (s^2/n)^2 / n instead of / (n - 1).
VERIFIED REPAIR
Use (va/na)^2/(na - 1) + (vb/nb)^2/(nb - 1).
Unsuccessful approach: Pooling the squared terms over na + nb - 2 is not the Satterthwaite form.
Case contract
a is control, b treatment. With sample variances (divisor n - 1), t = (mean_b - mean_a) / sqrt(va/na + vb/nb) and Welch-Satterthwaite df = (va/na + vb/nb)^2 / ((va/na)^2/(na-1) + (vb/nb)^2/(nb-1)). Fewer than two values in an arm or zero total variance -> None. Return [round(t, 6), round(df, 6)].
Why this case matters
Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa + sb) ** 2 / (sa ** 2 / na + sb ** 2 / nb)
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.713602, 7.953743] | [2.713602, 6.594671] | Failed |
| two-point arms | [2.683282, 2.941176] | [2.683282, 1.470588] | Failed |
| one arm constant | [0.866025, 3.0] | [0.866025, 2.0] | Failed |
| equal arms give zero t | [0.0, 6.0] | [0.0, 4.0] | Failed |
| negative effect | [-7.778175, 5.468354] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [1.290726, 9.302438] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [1.302114, 8.508486] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [2.622967, 4.443729] | [2.622967, 3.321892] | Failed |
SHA-256 / 3d5b4d10bad5ca65ae0e6e816ba775e33678f6ad2f43f669f47d97de47351c9b
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa + sb) ** 2 / ((sa ** 2 + sb ** 2) / (na + nb - 2))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.713602, 10.76885] | [2.713602, 6.594671] | Failed |
| two-point arms | [2.683282, 2.941176] | [2.683282, 1.470588] | Failed |
| one arm constant | [0.866025, 4.0] | [0.866025, 2.0] | Failed |
| equal arms give zero t | [0.0, 8.0] | [0.0, 4.0] | Failed |
| negative effect | [-7.778175, 6.923077] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [1.290726, 15.925698] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [1.302114, 13.613578] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [2.622967, 4.458366] | [2.622967, 3.321892] | Failed |
SHA-256 / 499489c06bc45ef98be8d625c57d2a3ec22b47fe13c8d1e2cdfb639cf0c73a94
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.713602, 6.594671] | [2.713602, 6.594671] | Passed |
| two-point arms | [2.683282, 1.470588] | [2.683282, 1.470588] | Passed |
| one arm constant | [0.866025, 2.0] | [0.866025, 2.0] | Passed |
| equal arms give zero t | [0.0, 4.0] | [0.0, 4.0] | Passed |
| negative effect | [-7.778175, 4.075472] | [-7.778175, 4.075472] | Passed |
| metric sample 1 | [1.290726, 7.219744] | [1.290726, 7.219744] | Passed |
| metric sample 2 | [1.302114, 6.806789] | [1.302114, 6.806789] | Passed |
| metric sample 3 | [2.622967, 3.321892] | [2.622967, 3.321892] | Passed |
SHA-256 / 54bef7dfe4d4551012597af385bda3f1c469a2e21e99d12426b213657335d6a6
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.520725+00:00.
Case digest / 82f239fd2d88b6e726ac57540231e94c2557fb1349abd0b61af16d016021e227