FA-74406 / Experiment statistics / Open access
Welch unequal-variance t statistic: Welch df numerator uses raw variances · case 01
Degrees of freedom scale with the metric's magnitude and sample size in nonsensical ways.
ROOT CAUSE
The numerator squares va + vb rather than va/na + vb/nb.
VERIFIED REPAIR
Square the sum of the per-arm variances of the mean.
Unsuccessful approach: Summing the squares instead of squaring the sum understates df.
Case contract
a is control, b treatment. With sample variances (divisor n - 1), t = (mean_b - mean_a) / sqrt(va/na + vb/nb) and Welch-Satterthwaite df = (va/na + vb/nb)^2 / ((va/na)^2/(na-1) + (vb/nb)^2/(nb-1)). Fewer than two values in an arm or zero total variance -> None. Return [round(t, 6), round(df, 6)].
Why this case matters
Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (va + vb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.713602, 214.033105] | [2.713602, 6.594671] | Failed |
| two-point arms | [2.683282, 5.882353] | [2.683282, 1.470588] | Failed |
| one arm constant | [0.866025, 18.0] | [0.866025, 2.0] | Failed |
| equal arms give zero t | [0.0, 36.0] | [0.0, 4.0] | Failed |
| negative effect | [-7.778175, 59.886792] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [1.290726, 175.595848] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [1.302114, 170.169719] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [2.622967, 50.30028] | [2.622967, 3.321892] | Failed |
SHA-256 / e6cdd66e37cff6a4678efa70970de5e9ec89e4682fe578c0b0dbdc9bc62329b3
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa ** 2 + sb ** 2) / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.713602, 4.899071] | [2.713602, 6.594671] | Failed |
| two-point arms | [2.683282, 1.0] | [2.683282, 1.470588] | Failed |
| one arm constant | [0.866025, 2.0] | [0.866025, 2.0] | Passed |
| equal arms give zero t | [0.0, 2.0] | [0.0, 4.0] | Failed |
| negative effect | [-7.778175, 2.943396] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [1.290726, 3.626714] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [1.302114, 4.0] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [2.622967, 2.980367] | [2.622967, 3.321892] | Failed |
SHA-256 / 15422d7b929696a776ab5890bc5977a312be0ab73d3349718a58c93104666a28
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892]),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 10', [[12, 13, 10, 4, 6], [29, 16, 23, 26, 11, 7, 18, 20, 10]], [2.882751, 11.994853]),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 24', [[18, 12, 5, 4, 0, 0, 6], [21, 22, 12, 13]], [2.940849, 7.679354]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.713602, 6.594671] | [2.713602, 6.594671] | Passed |
| two-point arms | [2.683282, 1.470588] | [2.683282, 1.470588] | Passed |
| one arm constant | [0.866025, 2.0] | [0.866025, 2.0] | Passed |
| equal arms give zero t | [0.0, 4.0] | [0.0, 4.0] | Passed |
| negative effect | [-7.778175, 4.075472] | [-7.778175, 4.075472] | Passed |
| metric sample 1 | [1.290726, 7.219744] | [1.290726, 7.219744] | Passed |
| metric sample 2 | [1.302114, 6.806789] | [1.302114, 6.806789] | Passed |
| metric sample 3 | [2.622967, 3.321892] | [2.622967, 3.321892] | Passed |
SHA-256 / 0047f8c5b478366121b9a70d496b67c2a2f83aac6329bde0f3d36872b87319e6
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.574777+00:00.
Case digest / 28b86439cdfc9aedb2411d421bbafa43571ce9bd42f2f63cfeffb3afbd378d52