FA-74401 / Experiment statistics / Open access
Welch unequal-variance t statistic: The t denominator mixes variances before dividing · case 01
Unequal arm sizes give a standard error that is neither pooled nor Welch.
ROOT CAUSE
The standard error is sqrt((va + vb) / (na + nb)).
THE FAILURE
The standard error is sqrt((va + vb) / (na + nb)).
Unsuccessful approach: Adding the two standard errors instead of the variances overstates uncertainty.
Case contract
a is control, b treatment. With sample variances (divisor n - 1), t = (mean_b - mean_a) / sqrt(va/na + vb/nb) and Welch-Satterthwaite df = (va/na + vb/nb)^2 / ((va/na)^2/(na-1) + (vb/nb)^2/(nb-1)). Fewer than two values in an arm or zero total variance -> None. Return [round(t, 6), round(df, 6)].
Why this case matters
Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / math.sqrt((va + vb) / (na + nb))
df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 5', [[15, 1, 9, 13, 9], [25, 29, 28, 1, 29]], [2.199917, 5.520904]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156]),
('metric sample 13', [[7, 13, 5, 4, 12], [6, 24, 6, 27, 4, 8, 3, 17, 7, 17]], [1.122702, 12.999998])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264]),
('metric sample 19', [[17, 1, 8, 10, 2], [22, 17]], [3.102706, 3.799185])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932]),
('metric sample 26', [[1, 11], [15, 23, 16]], [2.143769, 1.522005])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [3.59521, 6.594671] | [2.713602, 6.594671] | Failed |
| two-point arms | [3.794733, 1.470588] | [2.683282, 1.470588] | Failed |
| one arm constant | [1.224745, 2.0] | [0.866025, 2.0] | Failed |
| equal arms give zero t | [0.0, 4.0] | [0.0, 4.0] | Passed |
| negative effect | [-10.510864, 4.075472] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [1.837959, 7.219744] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [1.841468, 6.806789] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [3.257033, 3.321892] | [2.622967, 3.321892] | Failed |
SHA-256 / b2313cd6d9ad985b0f7e78098091d8978106baf73462af32a36269b1c31cd15f
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na < 2 or nb < 2:
return None
ma, mb = sum(a) / na, sum(b) / nb
va = sum((x - ma) ** 2 for x in a) / (na - 1)
vb = sum((x - mb) ** 2 for x in b) / (nb - 1)
sa, sb = va / na, vb / nb
if sa + sb == 0:
return None
t = (mb - ma) / (math.sqrt(sa) + math.sqrt(sb))
df = (sa + sb) ** 2 / (sa ** 2 / (na - 1) + sb ** 2 / (nb - 1))
return [round(t, 6), round(df, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('metric sample 1', [[0, 18, 6, 6], [3, 7, 11, 17, 25, 22]], [1.290726, 7.219744]),
('metric sample 2', [[12, 18, 10, 0, 10], [15, 23, 29, 2, 16]], [1.302114, 6.806789]),
('metric sample 3', [[5, 3], [21, 5, 12, 23]], [2.622967, 3.321892])],
[('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 4', [[6, 14, 12, 14, 7, 11, 3], [17, 29, 25, 22, 10]], [3.004687, 5.927117]),
('metric sample 5', [[15, 1, 9, 13, 9], [25, 29, 28, 1, 29]], [2.199917, 5.520904]),
('metric sample 6', [[13, 17, 16, 14, 0, 13], [10, 0, 6, 6, 0, 16, 5, 16]], [-1.428489, 10.999617])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 11', [[6, 10], [16, 12, 20, 6, 18]], [2.007859, 4.050223]),
('metric sample 12', [[7, 0], [27, 17, 26, 19, 20, 24, 1]], [3.236139, 3.199156]),
('metric sample 13', [[7, 13, 5, 4, 12], [6, 24, 6, 27, 4, 8, 3, 17, 7, 17]], [1.122702, 12.999998])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('equal arms give zero t', [[1, 2, 3], [1, 2, 3]], [0.0, 4.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 16',
[[13, 7, 4, 9, 13, 12], [20, 21, 17, 13, 14, 25, 13, 11, 16, 14, 6]],
[2.66816, 13.725637]),
('metric sample 18', [[18, 12, 1, 1], [12, 28, 6, 18, 8, 4, 15, 27, 24]], [1.497776, 6.245264]),
('metric sample 19', [[17, 1, 8, 10, 2], [22, 17]], [3.102706, 3.799185])],
[('unequal sizes and variances', [[1, 2, 3, 4], [2, 4, 6, 8, 10, 12]], [2.713602, 6.594671]),
('two-point arms', [[0, 2], [5, 9]], [2.683282, 1.470588]),
('one arm constant', [[3, 3, 3], [1, 5, 9]], [0.866025, 2.0]),
('negative effect', [[10, 12, 14, 16], [1, 3, 2]], [-7.778175, 4.075472]),
('single observation arm', [[1], [2, 3]], None),
('metric sample 21', [[0, 8, 15, 10], [28, 21, 11, 23, 28, 28]], [3.601238, 6.91205]),
('metric sample 25', [[17, 3], [0, 10, 29, 29, 25, 3, 18]], [0.751104, 1.981932]),
('metric sample 26', [[1, 11], [15, 23, 16]], [2.143769, 1.522005])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal sizes and variances | [2.070848, 6.594671] | [2.713602, 6.594671] | Failed |
| two-point arms | [2.0, 1.470588] | [2.683282, 1.470588] | Failed |
| one arm constant | [0.866025, 2.0] | [0.866025, 2.0] | Passed |
| equal arms give zero t | [0.0, 4.0] | [0.0, 4.0] | Passed |
| negative effect | [-5.887564, 4.075472] | [-7.778175, 4.075472] | Failed |
| metric sample 1 | [0.913214, 7.219744] | [1.290726, 7.219744] | Failed |
| metric sample 2 | [0.942638, 6.806789] | [1.302114, 6.806789] | Failed |
| metric sample 3 | [2.175666, 3.321892] | [2.622967, 3.321892] | Failed |
SHA-256 / 58e516253e88ce2083956442f8a42f67531097115a8cdbc53430a14c6200a9c8
HELD IN THE MEMBER ARCHIVE
The verified repair and its recorded checks are member-only.
This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.
Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.
Member access is invitation-based. Sign in with your invited account to inspect the repair.
Sign in to the archive ↗Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.560627+00:00.
Case digest / 3a666e7350655b81ceb7050afdb354b4219d2ba663bca956bc3d0b7cc60e6803