FAILURE MAP
← Case archive

FA-13986 / Numerical aggregation / Open access

Binary calibration decomposition: The global outcome rate is averaged over bin rates. · case 01

The reduction disagrees with its explicit aggregation oracle.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The global outcome rate is averaged over bin rates.

VERIFIED REPAIR

Preserve the binary calibration decomposition contract at the identified reduction decision.

Unsuccessful approach: Counting bins with any positive outcome also ignores observation frequency.

Case contract

Each [forecast numerator, positive denominator, positive outcome count, negative outcome count] describes binary observations with forecast in [0,1]. Merge identical rational forecasts. Return exact [reliability,resolution,uncertainty] Brier decomposition components. Empty total returns None; zero-size groups contribute nothing.

Why this case matters

Exact bounded examples isolate a reduction defect without floating-point or external-service effects.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
    groups={}
    for pn,pd,positive,negative in bins:
        p=Fraction(pn,pd)
        old=groups.get(p,(0,0))
        groups[p]=(old[0]+positive,old[1]+negative)
    n=sum(a+b for a,b in groups.values())
    if not n: return None
    base=sum((Fraction(a,a+b) for a,b in groups.values() if a+b),Fraction(0))/max(1,sum(a+b>0 for a,b in groups.values()))
    reliability=resolution=Fraction(0)
    for p,(a,b) in groups.items():
        k=a+b
        if not k: continue
        rate=Fraction(a,k)
        reliability+=Fraction(k,n)*(p-rate)**2
        resolution+=Fraction(k,n)*(rate-base)**2
    return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
regression 1['0', '1/16', '1/4']['0', '1/16', '1/4']Passed
regression 2['0', '1/16', '3/16']['0', '3/64', '15/64']Failed
regression 3NoneNonePassed
regression 4['0', '1/4', '1/4']['0', '1/4', '1/4']Passed
regression 5['16/225', '0', '6/25']['16/225', '0', '6/25']Passed
regression 6['23/75', '49/900', '221/900']['23/75', '49/960', '15/64']Failed
regression 7['0', '0', '0']['0', '0', '0']Passed
variable forecast count['0', '1/4', '1/4']['0', '1/4', '1/4']Passed

SHA-256 / edd3ca558a0c05aa3989d2a23fb8e5d887b5f1f075f14cb2f67c894b6217587f

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
    groups={}
    for pn,pd,positive,negative in bins:
        p=Fraction(pn,pd)
        old=groups.get(p,(0,0))
        groups[p]=(old[0]+positive,old[1]+negative)
    n=sum(a+b for a,b in groups.values())
    if not n: return None
    base=Fraction(sum(a>0 for a,b in groups.values()),max(1,len(groups)))
    reliability=resolution=Fraction(0)
    for p,(a,b) in groups.items():
        k=a+b
        if not k: continue
        rate=Fraction(a,k)
        reliability+=Fraction(k,n)*(p-rate)**2
        resolution+=Fraction(k,n)*(rate-base)**2
    return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
regression 1['0', '5/16', '0']['0', '1/16', '1/4']Failed
regression 2['0', '3/16', '0']['0', '3/64', '15/64']Failed
regression 3NoneNonePassed
regression 4['0', '1/4', '1/4']['0', '1/4', '1/4']Passed
regression 5['16/225', '1/100', '1/4']['16/225', '0', '6/25']Failed
regression 6['23/75', '23/120', '0']['23/75', '49/960', '15/64']Failed
regression 7['0', '0', '0']['0', '0', '0']Passed
variable forecast count['0', '1/4', '1/4']['0', '1/4', '1/4']Passed

SHA-256 / 1e9da053b31dc80a8a75858f898e89703f320c67bbc622e0a4435671e35ae3e0

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
    groups={}
    for pn,pd,positive,negative in bins:
        p=Fraction(pn,pd)
        old=groups.get(p,(0,0))
        groups[p]=(old[0]+positive,old[1]+negative)
    n=sum(a+b for a,b in groups.values())
    if not n: return None
    base=Fraction(sum(a for a,b in groups.values()),n)
    reliability=resolution=Fraction(0)
    for p,(a,b) in groups.items():
        k=a+b
        if not k: continue
        rate=Fraction(a,k)
        reliability+=Fraction(k,n)*(p-rate)**2
        resolution+=Fraction(k,n)*(rate-base)**2
    return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
regression 1['0', '1/16', '1/4']['0', '1/16', '1/4']Passed
regression 2['0', '3/64', '15/64']['0', '3/64', '15/64']Passed
regression 3NoneNonePassed
regression 4['0', '1/4', '1/4']['0', '1/4', '1/4']Passed
regression 5['16/225', '0', '6/25']['16/225', '0', '6/25']Passed
regression 6['23/75', '49/960', '15/64']['23/75', '49/960', '15/64']Passed
regression 7['0', '0', '0']['0', '0', '0']Passed
variable forecast count['0', '1/4', '1/4']['0', '1/4', '1/4']Passed

SHA-256 / 03d2899f3cf3a0d2d5f43bff6823c5ec9b99d0dd4fcdb428051c0b6e4f23c702

Verification & scope

Small offline integer/rational inputs only; no performance, statistical inference, or production-library conformance claim. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:39:12.358684+00:00.

Case digest / 22b6335e145602ab6e77483c431d8ba98a1464bcc0d34eddf3817c1d2f8b23d8