FA-74686 / Experiment statistics / Open access
Rank-sum test with ties: z is centred on the wrong null mean · case 01
Balanced data produces strongly negative z-scores.
ROOT CAUSE
The null expectation of U is taken as n_a n_b instead of n_a n_b / 2.
VERIFIED REPAIR
Centre U at n_a n_b / 2.
Unsuccessful approach: Centring at (n_a + n_b) / 2 confuses the U scale with the rank scale.
Case contract
Pool a (control) and b (treatment), assign mid-ranks to ties (1-based). U = R_b - n_b(n_b + 1)/2. Tie-corrected variance = n_a n_b / 12 * ((N + 1) - sum(t^3 - t) / (N(N - 1))); z = (U - n_a n_b / 2) / sqrt(variance) without continuity correction; zero variance gives z = 0. Empty arm -> None. Return [U, round(z, 6)].
Why this case matters
Rank tests are used for heavy-tailed metrics such as latency; ties are common in bucketed data.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na == 0 or nb == 0:
return None
pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
N = na + nb
ranks = [0.0] * N
ties = 0
i = 0
while i < N:
j = i
while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
j += 1
mid = (i + j) / 2 + 1
for k in range(i, j + 1):
ranks[k] = mid
t = j - i + 1
ties += t ** 3 - t
i = j + 1
rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
u = rb - nb * (nb + 1) / 2
var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
if var <= 0:
return [u, 0.0]
return [u, round((u - na * nb) / math.sqrt(var), 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
[('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282]),
('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934])],
[('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 9', [[5, 3, 4], [1, 1]], [0.0, -1.777047]),
('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
('rank sample 12', [[2, 1, 4], [6, 0, 3, 7]], [8.0, 0.707107])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
('rank sample 17', [[3, 5, 0], [0, 6, 2, 4, 4, 6]], [11.5, 0.65372]),
('rank sample 19', [[1, 3, 3, 0], [6, 5]], [8.0, 1.878673])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
('rank sample 27', [[2, 4, 4, 4], [1, 6]], [4.0, 0.0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ties get average ranks | [8.0, -0.463739] | [8.0, 1.623086] | Failed |
| two-way tie across arms | [3.5, -0.408248] | [3.5, 1.224745] | Failed |
| three-way tie with treatment | [3.0, -1.0] | [3.0, 1.0] | Failed |
| no ties | [9.0, -1.06066] | [9.0, 1.06066] | Failed |
| complete separation | [6.0, 0.0] | [6.0, 1.732051] | Failed |
| unequal arm sizes | [5.0, -1.475608] | [5.0, 0.491869] | Failed |
| rank sample 1 | [3.0, 0.0] | [3.0, 1.341641] | Failed |
| rank sample 2 | [18.0, -2.237127] | [18.0, 0.559282] | Failed |
SHA-256 / 7e923b176b94d8a361748035cdcfb258a41d9f3beddbf7c870bf60f06dbcf755
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na == 0 or nb == 0:
return None
pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
N = na + nb
ranks = [0.0] * N
ties = 0
i = 0
while i < N:
j = i
while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
j += 1
mid = (i + j) / 2 + 1
for k in range(i, j + 1):
ranks[k] = mid
t = j - i + 1
ties += t ** 3 - t
i = j + 1
rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
u = rb - nb * (nb + 1) / 2
var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
if var <= 0:
return [u, 0.0]
return [u, round((u - (na + nb) / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
[('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282]),
('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934])],
[('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 9', [[5, 3, 4], [1, 1]], [0.0, -1.777047]),
('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
('rank sample 12', [[2, 1, 4], [6, 0, 3, 7]], [8.0, 0.707107])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
('rank sample 17', [[3, 5, 0], [0, 6, 2, 4, 4, 6]], [11.5, 0.65372]),
('rank sample 19', [[1, 3, 3, 0], [6, 5]], [8.0, 1.878673])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
('rank sample 27', [[2, 4, 4, 4], [1, 6]], [4.0, 0.0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ties get average ranks | [8.0, 2.318694] | [8.0, 1.623086] | Failed |
| two-way tie across arms | [3.5, 1.224745] | [3.5, 1.224745] | Passed |
| three-way tie with treatment | [3.0, 1.0] | [3.0, 1.0] | Passed |
| no ties | [9.0, 1.944544] | [9.0, 1.06066] | Failed |
| complete separation | [6.0, 2.020726] | [6.0, 1.732051] | Failed |
| unequal arm sizes | [5.0, 0.983739] | [5.0, 0.491869] | Failed |
| rank sample 1 | [3.0, 0.894427] | [3.0, 1.341641] | Failed |
| rank sample 2 | [18.0, 2.330341] | [18.0, 0.559282] | Failed |
SHA-256 / 7bb44188fe4b2be1b6ab7c8cbb99873debd7a598fd53c13631461f7b38b7314c
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na == 0 or nb == 0:
return None
pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
N = na + nb
ranks = [0.0] * N
ties = 0
i = 0
while i < N:
j = i
while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
j += 1
mid = (i + j) / 2 + 1
for k in range(i, j + 1):
ranks[k] = mid
t = j - i + 1
ties += t ** 3 - t
i = j + 1
rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
u = rb - nb * (nb + 1) / 2
var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
if var <= 0:
return [u, 0.0]
return [u, round((u - na * nb / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
[('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282]),
('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934])],
[('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 9', [[5, 3, 4], [1, 1]], [0.0, -1.777047]),
('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
('rank sample 12', [[2, 1, 4], [6, 0, 3, 7]], [8.0, 0.707107])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
('rank sample 17', [[3, 5, 0], [0, 6, 2, 4, 4, 6]], [11.5, 0.65372]),
('rank sample 19', [[1, 3, 3, 0], [6, 5]], [8.0, 1.878673])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
('rank sample 27', [[2, 4, 4, 4], [1, 6]], [4.0, 0.0])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ties get average ranks | [8.0, 1.623086] | [8.0, 1.623086] | Passed |
| two-way tie across arms | [3.5, 1.224745] | [3.5, 1.224745] | Passed |
| three-way tie with treatment | [3.0, 1.0] | [3.0, 1.0] | Passed |
| no ties | [9.0, 1.06066] | [9.0, 1.06066] | Passed |
| complete separation | [6.0, 1.732051] | [6.0, 1.732051] | Passed |
| unequal arm sizes | [5.0, 0.491869] | [5.0, 0.491869] | Passed |
| rank sample 1 | [3.0, 1.341641] | [3.0, 1.341641] | Passed |
| rank sample 2 | [18.0, 0.559282] | [18.0, 0.559282] | Passed |
SHA-256 / 1cea135a2fa88212ff287542501569625f0fefd2e1b996188de0076dad859759
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:59.069992+00:00.
Case digest / eb55631bd83e367489ec76fcbb048546161d75621689ae77e33d20adc484144d