FA-74671 / Experiment statistics / Open access
Rank-sum test with ties: Tied values receive ordinal ranks · case 01
The statistic depends on arbitrary ordering of tied values and favours whichever arm sorts first.
ROOT CAUSE
Each tie group is given the rank of its first position.
VERIFIED REPAIR
Assign every member of a tie group the average of its positions.
Unsuccessful approach: Integer division truncates half-integer midranks.
Case contract
Pool a (control) and b (treatment), assign mid-ranks to ties (1-based). U = R_b - n_b(n_b + 1)/2. Tie-corrected variance = n_a n_b / 12 * ((N + 1) - sum(t^3 - t) / (N(N - 1))); z = (U - n_a n_b / 2) / sqrt(variance) without continuity correction; zero variance gives z = 0. Empty arm -> None. Return [U, round(z, 6)].
Why this case matters
Rank tests are used for heavy-tailed metrics such as latency; ties are common in bucketed data.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na == 0 or nb == 0:
return None
pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
N = na + nb
ranks = [0.0] * N
ties = 0
i = 0
while i < N:
j = i
while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
j += 1
mid = i + 1
for k in range(i, j + 1):
ranks[k] = mid
t = j - i + 1
ties += t ** 3 - t
i = j + 1
rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
u = rb - nb * (nb + 1) / 2
var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
if var <= 0:
return [u, 0.0]
return [u, round((u - na * nb / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
[('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109])],
[('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
('rank sample 14', [[5], [1, 1, 4, 7]], [1.0, -0.725476]),
('rank sample 28', [[0, 4, 0, 0, 3], [6, 4, 5]], [14.5, 2.152028])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
('rank sample 37', [[1, 3, 3], [4, 6, 2, 1, 2, 2]], [9.5, 0.132453])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
('rank sample 32', [[3, 2, 4, 5, 1], [6, 3, 3, 7, 6, 3]], [22.5, 1.404879]),
('rank sample 50', [[0, 2, 2, 2, 4], [4]], [4.5, 1.264911])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ties get average ranks | [7.0, 1.159347] | [8.0, 1.623086] | Failed |
| two-way tie across arms | [3.0, 0.816497] | [3.5, 1.224745] | Failed |
| three-way tie with treatment | [2.0, 0.0] | [3.0, 1.0] | Failed |
| no ties | [9.0, 1.06066] | [9.0, 1.06066] | Passed |
| complete separation | [6.0, 1.732051] | [6.0, 1.732051] | Passed |
| unequal arm sizes | [4.0, 0.0] | [5.0, 0.491869] | Failed |
| rank sample 1 | [3.0, 1.341641] | [3.0, 1.341641] | Passed |
| rank sample 2 | [14.0, -0.186427] | [18.0, 0.559282] | Failed |
SHA-256 / 86b1b3a03ce95c17bc2635cff1226c735f3c00a5b3010aaca0aebe77d94cab09
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na == 0 or nb == 0:
return None
pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
N = na + nb
ranks = [0.0] * N
ties = 0
i = 0
while i < N:
j = i
while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
j += 1
mid = (i + j) // 2 + 1
for k in range(i, j + 1):
ranks[k] = mid
t = j - i + 1
ties += t ** 3 - t
i = j + 1
rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
u = rb - nb * (nb + 1) / 2
var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
if var <= 0:
return [u, 0.0]
return [u, round((u - na * nb / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
[('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109])],
[('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
('rank sample 14', [[5], [1, 1, 4, 7]], [1.0, -0.725476]),
('rank sample 28', [[0, 4, 0, 0, 3], [6, 4, 5]], [14.5, 2.152028])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
('rank sample 37', [[1, 3, 3], [4, 6, 2, 1, 2, 2]], [9.5, 0.132453])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
('rank sample 32', [[3, 2, 4, 5, 1], [6, 3, 3, 7, 6, 3]], [22.5, 1.404879]),
('rank sample 50', [[0, 2, 2, 2, 4], [4]], [4.5, 1.264911])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ties get average ranks | [8.0, 1.623086] | [8.0, 1.623086] | Passed |
| two-way tie across arms | [3.0, 0.816497] | [3.5, 1.224745] | Failed |
| three-way tie with treatment | [3.0, 1.0] | [3.0, 1.0] | Passed |
| no ties | [9.0, 1.06066] | [9.0, 1.06066] | Passed |
| complete separation | [6.0, 1.732051] | [6.0, 1.732051] | Passed |
| unequal arm sizes | [5.0, 0.491869] | [5.0, 0.491869] | Passed |
| rank sample 1 | [3.0, 1.341641] | [3.0, 1.341641] | Passed |
| rank sample 2 | [18.0, 0.559282] | [18.0, 0.559282] | Passed |
SHA-256 / ec4d40e6e9504a38bb4610b8856018df2c65134308bb09929f5d1f6466078440
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
na, nb = len(a), len(b)
if na == 0 or nb == 0:
return None
pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
N = na + nb
ranks = [0.0] * N
ties = 0
i = 0
while i < N:
j = i
while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
j += 1
mid = (i + j) / 2 + 1
for k in range(i, j + 1):
ranks[k] = mid
t = j - i + 1
ties += t ** 3 - t
i = j + 1
rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
u = rb - nb * (nb + 1) / 2
var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
if var <= 0:
return [u, 0.0]
return [u, round((u - na * nb / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
[('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109])],
[('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
('rank sample 14', [[5], [1, 1, 4, 7]], [1.0, -0.725476]),
('rank sample 28', [[0, 4, 0, 0, 3], [6, 4, 5]], [14.5, 2.152028])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
('rank sample 37', [[1, 3, 3], [4, 6, 2, 1, 2, 2]], [9.5, 0.132453])],
[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
('rank sample 32', [[3, 2, 4, 5, 1], [6, 3, 3, 7, 6, 3]], [22.5, 1.404879]),
('rank sample 50', [[0, 2, 2, 2, 4], [4]], [4.5, 1.264911])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| ties get average ranks | [8.0, 1.623086] | [8.0, 1.623086] | Passed |
| two-way tie across arms | [3.5, 1.224745] | [3.5, 1.224745] | Passed |
| three-way tie with treatment | [3.0, 1.0] | [3.0, 1.0] | Passed |
| no ties | [9.0, 1.06066] | [9.0, 1.06066] | Passed |
| complete separation | [6.0, 1.732051] | [6.0, 1.732051] | Passed |
| unequal arm sizes | [5.0, 0.491869] | [5.0, 0.491869] | Passed |
| rank sample 1 | [3.0, 1.341641] | [3.0, 1.341641] | Passed |
| rank sample 2 | [18.0, 0.559282] | [18.0, 0.559282] | Passed |
SHA-256 / 3254fff0a031bc03bfea6577d5e2896295f7a4814d31136102e065a4511583d0
Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.986513+00:00.
Case digest / eed79e57cb8e4a7f1998d9b3bf7736b6fe8838b2dfcd9180da32f2ca30f290b4