FAILURE MAP
← Case archive

FA-11866 / Search retrieval semantics / Open access

Rank fusion uses incomparable raw scores · case 01

Rank fusion uses incomparable raw scores.

Verified by executionVariant 1 · 6 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

Scores from independent retrievers are summed despite incompatible scales.

THE FAILURE

Scores from independent retrievers are summed despite incompatible scales.

Unsuccessful approach: Reciprocal zero-based ranks shift every contribution and change cross-list tradeoffs.

Case contract

Each input list contains unique [id,score] hits in ranked order. Fuse by sum(1/(offset+rank)) using one-based rank; return IDs descending by exact rational fused score, lexical ties. Offset is a positive integer.

Why this case matters

An offline deterministic retrieval model isolates this search contract from tokenization, storage, and network behavior.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(lists, offset):
    scores={}
    for hits in lists:
        for ident,score in hits: scores[ident]=scores.get(ident,0)+score
    return sorted(scores,key=lambda ident:(-scores[ident],ident))
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('scale irrelevant', solve([[['a',N],['b',1000*N]],[['b',N]]],1),['b','a'])
check('one versus two deep ranks', solve([[['a',100],['x',2],['b',1]],[['y',5],['z',4],['b',1]]],1),['a','b','y','x','z'])
check('lexical ties', solve([[['z',N]],[['a',2*N]]],2),['a','z'])
check('empty lists', solve([[],[]],N),[])
check('single list rank order', solve([[['z',1],['a',100]]],N),['z','a'])
check('same document two lists', solve([[['a',0]],[['a',0]]],N),['a'])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
scale irrelevant['b', 'a']['b', 'a']Passed
one versus two deep ranks['a', 'y', 'z', 'b', 'x']['a', 'b', 'y', 'x', 'z']Failed
lexical ties['a', 'z']['a', 'z']Passed
empty lists[][]Passed
single list rank order['a', 'z']['z', 'a']Failed
same document two lists['a']['a']Passed

SHA-256 / 49968d997efef2c4cf9608d328dc709bd307cdbde2a5802fd03dccaa0dd3b316

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(lists, offset):
    from fractions import Fraction
    scores={}
    for hits in lists:
        for rank,(ident,score) in enumerate(hits): scores[ident]=scores.get(ident,0)+Fraction(1,offset+rank)
    return sorted(scores,key=lambda ident:(-scores[ident],ident))
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('scale irrelevant', solve([[['a',N],['b',1000*N]],[['b',N]]],1),['b','a'])
check('one versus two deep ranks', solve([[['a',100],['x',2],['b',1]],[['y',5],['z',4],['b',1]]],1),['a','b','y','x','z'])
check('lexical ties', solve([[['z',N]],[['a',2*N]]],2),['a','z'])
check('empty lists', solve([[],[]],N),[])
check('single list rank order', solve([[['z',1],['a',100]]],N),['z','a'])
check('same document two lists', solve([[['a',0]],[['a',0]]],N),['a'])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
scale irrelevant['b', 'a']['b', 'a']Passed
one versus two deep ranks['a', 'y', 'b', 'x', 'z']['a', 'b', 'y', 'x', 'z']Failed
lexical ties['a', 'z']['a', 'z']Passed
empty lists[][]Passed
single list rank order['z', 'a']['z', 'a']Passed
same document two lists['a']['a']Passed

SHA-256 / 7732f4e5fc7df32a05d752bd7db97e589081cfb30c562a02ee1c24240a29995c

HELD IN THE MEMBER ARCHIVE

The verified repair and its recorded checks are member-only.

This mechanism has 6 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.

Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.

Member access is invitation-based. Sign in with your invited account to inspect the repair.

Sign in to the archive ↗

Verification & scope

Inputs are already tokenized or scored; this model makes no claim about production engine performance or linguistic analysis. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:38:51.719578+00:00.

Case digest / b5d972334785f670fcc5293b7164b96b6eee030d7f919390c96725d471cc7eea