FAILURE MAP
← Case archive

FA-74671 / Experiment statistics / Open access

Rank-sum test with ties: Tied values receive ordinal ranks · case 01

The statistic depends on arbitrary ordering of tied values and favours whichever arm sorts first.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

Each tie group is given the rank of its first position.

VERIFIED REPAIR

Assign every member of a tie group the average of its positions.

Unsuccessful approach: Integer division truncates half-integer midranks.

Case contract

Pool a (control) and b (treatment), assign mid-ranks to ties (1-based). U = R_b - n_b(n_b + 1)/2. Tie-corrected variance = n_a n_b / 12 * ((N + 1) - sum(t^3 - t) / (N(N - 1))); z = (U - n_a n_b / 2) / sqrt(variance) without continuity correction; zero variance gives z = 0. Empty arm -> None. Return [U, round(z, 6)].

Why this case matters

Rank tests are used for heavy-tailed metrics such as latency; ties are common in bucketed data.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na == 0 or nb == 0:
        return None
    pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
    N = na + nb
    ranks = [0.0] * N
    ties = 0
    i = 0
    while i < N:
        j = i
        while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
            j += 1
        mid = i + 1
        for k in range(i, j + 1):
            ranks[k] = mid
        t = j - i + 1
        ties += t ** 3 - t
        i = j + 1
    rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
    u = rb - nb * (nb + 1) / 2
    var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
    if var <= 0:
        return [u, 0.0]
    return [u, round((u - na * nb / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
  ('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
 [('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934]),
  ('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109])],
 [('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
  ('rank sample 14', [[5], [1, 1, 4, 7]], [1.0, -0.725476]),
  ('rank sample 28', [[0, 4, 0, 0, 3], [6, 4, 5]], [14.5, 2.152028])],
 [('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
  ('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
  ('rank sample 37', [[1, 3, 3], [4, 6, 2, 1, 2, 2]], [9.5, 0.132453])],
 [('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
  ('rank sample 32', [[3, 2, 4, 5, 1], [6, 3, 3, 7, 6, 3]], [22.5, 1.404879]),
  ('rank sample 50', [[0, 2, 2, 2, 4], [4]], [4.5, 1.264911])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ties get average ranks[7.0, 1.159347][8.0, 1.623086]Failed
two-way tie across arms[3.0, 0.816497][3.5, 1.224745]Failed
three-way tie with treatment[2.0, 0.0][3.0, 1.0]Failed
no ties[9.0, 1.06066][9.0, 1.06066]Passed
complete separation[6.0, 1.732051][6.0, 1.732051]Passed
unequal arm sizes[4.0, 0.0][5.0, 0.491869]Failed
rank sample 1[3.0, 1.341641][3.0, 1.341641]Passed
rank sample 2[14.0, -0.186427][18.0, 0.559282]Failed

SHA-256 / 86b1b3a03ce95c17bc2635cff1226c735f3c00a5b3010aaca0aebe77d94cab09

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na == 0 or nb == 0:
        return None
    pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
    N = na + nb
    ranks = [0.0] * N
    ties = 0
    i = 0
    while i < N:
        j = i
        while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
            j += 1
        mid = (i + j) // 2 + 1
        for k in range(i, j + 1):
            ranks[k] = mid
        t = j - i + 1
        ties += t ** 3 - t
        i = j + 1
    rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
    u = rb - nb * (nb + 1) / 2
    var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
    if var <= 0:
        return [u, 0.0]
    return [u, round((u - na * nb / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
  ('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
 [('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934]),
  ('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109])],
 [('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
  ('rank sample 14', [[5], [1, 1, 4, 7]], [1.0, -0.725476]),
  ('rank sample 28', [[0, 4, 0, 0, 3], [6, 4, 5]], [14.5, 2.152028])],
 [('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
  ('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
  ('rank sample 37', [[1, 3, 3], [4, 6, 2, 1, 2, 2]], [9.5, 0.132453])],
 [('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
  ('rank sample 32', [[3, 2, 4, 5, 1], [6, 3, 3, 7, 6, 3]], [22.5, 1.404879]),
  ('rank sample 50', [[0, 2, 2, 2, 4], [4]], [4.5, 1.264911])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ties get average ranks[8.0, 1.623086][8.0, 1.623086]Passed
two-way tie across arms[3.0, 0.816497][3.5, 1.224745]Failed
three-way tie with treatment[3.0, 1.0][3.0, 1.0]Passed
no ties[9.0, 1.06066][9.0, 1.06066]Passed
complete separation[6.0, 1.732051][6.0, 1.732051]Passed
unequal arm sizes[5.0, 0.491869][5.0, 0.491869]Passed
rank sample 1[3.0, 1.341641][3.0, 1.341641]Passed
rank sample 2[18.0, 0.559282][18.0, 0.559282]Passed

SHA-256 / ec4d40e6e9504a38bb4610b8856018df2c65134308bb09929f5d1f6466078440

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(a, b):
    na, nb = len(a), len(b)
    if na == 0 or nb == 0:
        return None
    pooled = sorted([(v, 0) for v in a] + [(v, 1) for v in b])
    N = na + nb
    ranks = [0.0] * N
    ties = 0
    i = 0
    while i < N:
        j = i
        while j + 1 < N and pooled[j + 1][0] == pooled[i][0]:
            j += 1
        mid = (i + j) / 2 + 1
        for k in range(i, j + 1):
            ranks[k] = mid
        t = j - i + 1
        ties += t ** 3 - t
        i = j + 1
    rb = sum(r for r, (v, g) in zip(ranks, pooled) if g == 1)
    u = rb - nb * (nb + 1) / 2
    var = na * nb / 12 * ((N + 1) - ties / (N * (N - 1))) if N > 1 else 0.0
    if var <= 0:
        return [u, 0.0]
    return [u, round((u - na * nb / 2) / math.sqrt(var), 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 1', [[1, 3, 2], [7]], [3.0, 1.341641]),
  ('rank sample 2', [[3, 1, 2, 2, 4], [7, 3, 1, 1, 5, 3]], [18.0, 0.559282])],
 [('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 4', [[4], [3, 4, 2, 7, 2, 3]], [1.5, -0.770934]),
  ('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109])],
 [('three-way tie with treatment', [[5, 5], [5, 6]], [3.0, 1.0]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 11', [[1, 1, 2, 3, 4], [2]], [2.5, 0.0]),
  ('rank sample 14', [[5], [1, 1, 4, 7]], [1.0, -0.725476]),
  ('rank sample 28', [[0, 4, 0, 0, 3], [6, 4, 5]], [14.5, 2.152028])],
 [('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('no ties', [[1, 3, 5], [2, 4, 6, 8]], [9.0, 1.06066]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 16', [[4, 0, 3, 3, 1], [4, 1]], [6.0, 0.398109]),
  ('rank sample 23', [[3, 4], [0, 4, 6]], [3.5, 0.296174]),
  ('rank sample 37', [[1, 3, 3], [4, 6, 2, 1, 2, 2]], [9.5, 0.132453])],
 [('ties get average ranks', [[1, 2, 2], [2, 3, 4]], [8.0, 1.623086]),
  ('two-way tie across arms', [[1, 2], [2, 3]], [3.5, 1.224745]),
  ('complete separation', [[1, 2], [3, 4, 5]], [6.0, 1.732051]),
  ('all values tied', [[2, 2], [2, 2, 2]], [3.0, 0.0]),
  ('unequal arm sizes', [[0, 1, 1, 4], [1, 2]], [5.0, 0.491869]),
  ('rank sample 21', [[4, 4, 3, 1, 2], [4, 7, 7, 7, 6, 5]], [29.0, 2.603819]),
  ('rank sample 32', [[3, 2, 4, 5, 1], [6, 3, 3, 7, 6, 3]], [22.5, 1.404879]),
  ('rank sample 50', [[0, 2, 2, 2, 4], [4]], [4.5, 1.264911])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
ties get average ranks[8.0, 1.623086][8.0, 1.623086]Passed
two-way tie across arms[3.5, 1.224745][3.5, 1.224745]Passed
three-way tie with treatment[3.0, 1.0][3.0, 1.0]Passed
no ties[9.0, 1.06066][9.0, 1.06066]Passed
complete separation[6.0, 1.732051][6.0, 1.732051]Passed
unequal arm sizes[5.0, 0.491869][5.0, 0.491869]Passed
rank sample 1[3.0, 1.341641][3.0, 1.341641]Passed
rank sample 2[18.0, 0.559282][18.0, 0.559282]Passed

SHA-256 / 3254fff0a031bc03bfea6577d5e2896295f7a4814d31136102e065a4511583d0

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:58.986513+00:00.

Case digest / eed79e57cb8e4a7f1998d9b3bf7736b6fe8838b2dfcd9180da32f2ca30f290b4