FAILURE MAP
← Case archive

FA-13961 / Numerical aggregation / Open access

Binary calibration decomposition: Calibration components give every forecast bin equal weight. · case 01

The reduction disagrees with its explicit aggregation oracle.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

Calibration components give every forecast bin equal weight.

VERIFIED REPAIR

Preserve the binary calibration decomposition contract at the identified reduction decision.

Unsuccessful approach: Dividing by bin count instead of observation count does not normalize frequency mass.

Case contract

Each [forecast numerator, positive denominator, positive outcome count, negative outcome count] describes binary observations with forecast in [0,1]. Merge identical rational forecasts. Return exact [reliability,resolution,uncertainty] Brier decomposition components. Empty total returns None; zero-size groups contribute nothing.

Why this case matters

Exact bounded examples isolate a reduction defect without floating-point or external-service effects.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
    groups={}
    for pn,pd,positive,negative in bins:
        p=Fraction(pn,pd)
        old=groups.get(p,(0,0))
        groups[p]=(old[0]+positive,old[1]+negative)
    n=sum(a+b for a,b in groups.values())
    if not n: return None
    base=Fraction(sum(a for a,b in groups.values()),n)
    reliability=resolution=Fraction(0)
    for p,(a,b) in groups.items():
        k=a+b
        if not k: continue
        rate=Fraction(a,k)
        reliability+=Fraction(1,len(groups))*(p-rate)**2
        resolution+=Fraction(1,len(groups))*(rate-base)**2
    return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
regression 1['0', '1/16', '1/4']['0', '1/16', '1/4']Passed
regression 2['0', '5/64', '15/64']['0', '3/64', '15/64']Failed
regression 3NoneNonePassed
regression 4['0', '1/4', '1/4']['0', '1/4', '1/4']Passed
regression 5['8/225', '0', '6/25']['16/225', '0', '6/25']Failed
regression 6['13/45', '833/14400', '15/64']['23/75', '49/960', '15/64']Failed
regression 7['0', '0', '0']['0', '0', '0']Passed
variable forecast count['0', '1/4', '1/4']['0', '1/4', '1/4']Passed

SHA-256 / 60ea97960616b3ef10ebb86e9ac83613e347f25d85777de65cc2fb3f64057775

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
    groups={}
    for pn,pd,positive,negative in bins:
        p=Fraction(pn,pd)
        old=groups.get(p,(0,0))
        groups[p]=(old[0]+positive,old[1]+negative)
    n=sum(a+b for a,b in groups.values())
    if not n: return None
    base=Fraction(sum(a for a,b in groups.values()),n)
    reliability=resolution=Fraction(0)
    for p,(a,b) in groups.items():
        k=a+b
        if not k: continue
        rate=Fraction(a,k)
        reliability+=Fraction(k,len(groups))*(p-rate)**2
        resolution+=Fraction(k,len(groups))*(rate-base)**2
    return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
regression 1['0', '1/4', '1/4']['0', '1/16', '1/4']Failed
regression 2['0', '3/16', '15/64']['0', '3/64', '15/64']Failed
regression 3NoneNonePassed
regression 4['0', '1', '1/4']['0', '1/4', '1/4']Failed
regression 5['8/45', '0', '6/25']['16/225', '0', '6/25']Failed
regression 6['92/75', '49/240', '15/64']['23/75', '49/960', '15/64']Failed
regression 7['0', '0', '0']['0', '0', '0']Passed
variable forecast count['0', '1/4', '1/4']['0', '1/4', '1/4']Passed

SHA-256 / c945a5a20a5ce86c66d8888a7f3df45b4c46c208a5bf2059af5e00cefde4b912

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
    groups={}
    for pn,pd,positive,negative in bins:
        p=Fraction(pn,pd)
        old=groups.get(p,(0,0))
        groups[p]=(old[0]+positive,old[1]+negative)
    n=sum(a+b for a,b in groups.values())
    if not n: return None
    base=Fraction(sum(a for a,b in groups.values()),n)
    reliability=resolution=Fraction(0)
    for p,(a,b) in groups.items():
        k=a+b
        if not k: continue
        rate=Fraction(a,k)
        reliability+=Fraction(k,n)*(p-rate)**2
        resolution+=Fraction(k,n)*(rate-base)**2
    return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
regression 1['0', '1/16', '1/4']['0', '1/16', '1/4']Passed
regression 2['0', '3/64', '15/64']['0', '3/64', '15/64']Passed
regression 3NoneNonePassed
regression 4['0', '1/4', '1/4']['0', '1/4', '1/4']Passed
regression 5['16/225', '0', '6/25']['16/225', '0', '6/25']Passed
regression 6['23/75', '49/960', '15/64']['23/75', '49/960', '15/64']Passed
regression 7['0', '0', '0']['0', '0', '0']Passed
variable forecast count['0', '1/4', '1/4']['0', '1/4', '1/4']Passed

SHA-256 / 03d2899f3cf3a0d2d5f43bff6823c5ec9b99d0dd4fcdb428051c0b6e4f23c702

Verification & scope

Small offline integer/rational inputs only; no performance, statistical inference, or production-library conformance claim. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:39:12.269193+00:00.

Case digest / a529df96089bff94e654c932b04aa329d894322106e4d41d8cf1b06688c64153