FA-74391 / Experiment statistics / Open access
Welch unequal-variance t statistic: Arm variances use the population divisor · case 01
Small arms get understated variance and overconfident t statistics.
ROOT CAUSE
Variances divide by n instead of n - 1.
VERIFIED REPAIR
Divide the squared deviations by n - 1 in each arm.
Unsuccessful approach: Dividing by the pooled na + nb - 2 mixes a pooled-variance idea into Welch.
Case contract
a is control, b treatment. With sample variances (divisor n - 1), t = (mean_b - mean_a) / sqrt(va/na + vb/nb) and Welch-Satterthwaite df = (va/na + vb/nb)^2 / ((va/na)^2/(na-1) + (vb/nb)^2/(nb-1)). Fewer than two values in an arm or zero total variance -> None. Return [round(t, 6), round(df, 6)].
Why this case matters
Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / na
vb = sum((x - mb) ** 2 for x in b) / nb
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156]),
('metric sample 13', [[7, 13, 5, 4, 12], [6, 24, 6, 27, 4, 8, 3, 17, 7, 17]], [1.122702, 12.999998])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.995381, 6.45827] | [2.713602, 6.594671] | Failed |
| two-point arms | [3.794733, 1.470588] | [2.683282, 1.470588] | Failed |
| one arm constant | [1.06066, 2.0] | [0.866025, 2.0] | Failed |
| equal arms give zero t | [0.0, 4.0] | [0.0, 4.0] | Passed |
| negative effect | [-9.065797, 3.973126] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [1.453265, 7.439647] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [1.455808, 6.806789] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [3.05656, 3.220158] | [2.622967, 3.321892] | Failed |
SHA-256 / 7e8bb67bb68a6378d5fbdc418be8e23ed417f00180344d3873a6602f5332e01a
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na + nb - 2)
vb = sum((x - mb) ** 2 for x in b) / (na + nb - 2)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156]),
('metric sample 13', [[7, 13, 5, 4, 12], [6, 24, 6, 27, 4, 8, 3, 17, 7, 17]], [1.122702, 12.999998])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [3.54146, 6.013767] | [2.713602, 6.594671] | Failed |
| two-point arms | [3.794733, 1.470588] | [2.683282, 1.470588] | Failed |
| one arm constant | [1.224745, 2.0] | [0.866025, 2.0] | Failed |
| equal arms give zero t | [0.0, 4.0] | [0.0, 4.0] | Passed |
| negative effect | [-10.332701, 3.753247] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [1.841149, 7.963945] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [1.841468, 6.806789] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [3.08516, 3.112643] | [2.622967, 3.321892] | Failed |
SHA-256 / 0c3acea939dbb4b1dd0c396628eb9d2864bd20dd8e6c3724add6ac397e266a71
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt(sa + sb)
df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617]),
('metric sample 7', [[0, 9, 16, 8, 15, 5, 3], [29, 4, 20, 26, 11, 20]], [2.332551, 8.238859])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156]),
('metric sample 13', [[7, 13, 5, 4, 12], [6, 24, 6, 27, 4, 8, 3, 17, 7, 17]], [1.122702, 12.999998])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 17', [[10, 4, 0, 14], [6, 17, 1, 13, 11]], [0.622825, 6.572986]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 22', [[1, 5, 15], [29, 28, 6, 25, 5, 12]], [1.705196, 6.118881]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.713602, 6.594671] | [2.713602, 6.594671] | Passed |
| two-point arms | [2.683282, 1.470588] | [2.683282, 1.470588] | Passed |
| one arm constant | [0.866025, 2.0] | [0.866025, 2.0] | Passed |
| equal arms give zero t | [0.0, 4.0] | [0.0, 4.0] | Passed |
| negative effect | [-7.778175, 4.075472] | [-7.778175, 4.075472] | Passed |
| metric sample 1 | [1.290726, 7.219744] | [1.290726, 7.219744] | Passed |
| metric sample 2 | [1.302114, 6.806789] | [1.302114, 6.806789] | Passed |
| metric sample 3 | [2.622967, 3.321892] | [2.622967, 3.321892] | Passed |
SHA-256 / 55cc5b2be0520abe67645ab700fa8a85381675fe15916f10152acceea5261f4d
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.475596+00:00.
Case digest / f0e10d481af07c29a6f04bfe1c0e8045aa74763493f77a44896d9cdeb4c29474