FA-13966 / Numerical aggregation / Open access
Binary calibration decomposition: Reliability is measured around the global outcome rate. · case 01
The reduction disagrees with its explicit aggregation oracle.
ROOT CAUSE
Reliability is measured around the global outcome rate.
VERIFIED REPAIR
Preserve the binary calibration decomposition contract at the identified reduction decision.
Unsuccessful approach: Substituting resolution for reliability ignores the issued forecast.
Case contract
Each [forecast numerator, positive denominator, positive outcome count, negative outcome count] describes binary observations with forecast in [0,1]. Merge identical rational forecasts. Return exact [reliability,resolution,uncertainty] Brier decomposition components. Empty total returns None; zero-size groups contribute nothing.
Why this case matters
Exact bounded examples isolate a reduction defect without floating-point or external-service effects.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
groups={}
for pn,pd,positive,negative in bins:
p=Fraction(pn,pd)
old=groups.get(p,(0,0))
groups[p]=(old[0]+positive,old[1]+negative)
n=sum(a+b for a,b in groups.values())
if not n: return None
base=Fraction(sum(a for a,b in groups.values()),n)
reliability=resolution=Fraction(0)
for p,(a,b) in groups.items():
k=a+b
if not k: continue
rate=Fraction(a,k)
reliability+=Fraction(k,n)*(p-base)**2
resolution+=Fraction(k,n)*(rate-base)**2
return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| regression 1 | ['1/16', '1/16', '1/4'] | ['0', '1/16', '1/4'] | Failed |
| regression 2 | ['3/64', '3/64', '15/64'] | ['0', '3/64', '15/64'] | Failed |
| regression 3 | None | None | Passed |
| regression 4 | ['1/4', '1/4', '1/4'] | ['0', '1/4', '1/4'] | Failed |
| regression 5 | ['16/225', '0', '6/25'] | ['16/225', '0', '6/25'] | Passed |
| regression 6 | ['199/1600', '49/960', '15/64'] | ['23/75', '49/960', '15/64'] | Failed |
| regression 7 | ['0', '0', '0'] | ['0', '0', '0'] | Passed |
| variable forecast count | ['1/4', '1/4', '1/4'] | ['0', '1/4', '1/4'] | Failed |
SHA-256 / 1537159341d92f9d3d54d69803c8534f09ada9e65808f06f2833858604452f2e
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
groups={}
for pn,pd,positive,negative in bins:
p=Fraction(pn,pd)
old=groups.get(p,(0,0))
groups[p]=(old[0]+positive,old[1]+negative)
n=sum(a+b for a,b in groups.values())
if not n: return None
base=Fraction(sum(a for a,b in groups.values()),n)
reliability=resolution=Fraction(0)
for p,(a,b) in groups.items():
k=a+b
if not k: continue
rate=Fraction(a,k)
reliability+=Fraction(k,n)*(rate-base)**2
resolution+=Fraction(k,n)*(rate-base)**2
return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| regression 1 | ['1/16', '1/16', '1/4'] | ['0', '1/16', '1/4'] | Failed |
| regression 2 | ['3/64', '3/64', '15/64'] | ['0', '3/64', '15/64'] | Failed |
| regression 3 | None | None | Passed |
| regression 4 | ['1/4', '1/4', '1/4'] | ['0', '1/4', '1/4'] | Failed |
| regression 5 | ['0', '0', '6/25'] | ['16/225', '0', '6/25'] | Failed |
| regression 6 | ['49/960', '49/960', '15/64'] | ['23/75', '49/960', '15/64'] | Failed |
| regression 7 | ['0', '0', '0'] | ['0', '0', '0'] | Passed |
| variable forecast count | ['1/4', '1/4', '1/4'] | ['0', '1/4', '1/4'] | Failed |
SHA-256 / bca93081a7e05bd6ea37dfa39277b88897d255b97515360b1fd7e3d701964662
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
from collections import Counter, defaultdict
import math
import itertools
N = 1
observations = []
def solve(bins):
groups={}
for pn,pd,positive,negative in bins:
p=Fraction(pn,pd)
old=groups.get(p,(0,0))
groups[p]=(old[0]+positive,old[1]+negative)
n=sum(a+b for a,b in groups.values())
if not n: return None
base=Fraction(sum(a for a,b in groups.values()),n)
reliability=resolution=Fraction(0)
for p,(a,b) in groups.items():
k=a+b
if not k: continue
rate=Fraction(a,k)
reliability+=Fraction(k,n)*(p-rate)**2
resolution+=Fraction(k,n)*(rate-base)**2
return [str(reliability),str(resolution),str(base*(1-base))]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('regression 1', solve(*([(1, 4, 1, 3), (3, 4, 3, 1)],)), ['0', '1/16', '1/4'])
check('regression 2', solve(*([(1, 2, 3, 1), (1, 2, 0, 2), (1, 1, 2, 0)],)), ['0', '3/64', '15/64'])
check('regression 3', solve(*([],)), None)
check('regression 4', solve(*([(0, 1, 0, 4), (1, 1, 4, 0)],)), ['0', '1/4', '1/4'])
check('regression 5', solve(*([(1, 3, 0, 0), (2, 3, 2, 3)],)), ['16/225', '0', '6/25'])
check('regression 6', solve(*([(1, 5, 4, 1), (4, 5, 1, 2)],)), ['23/75', '49/960', '15/64'])
check('regression 7', solve(*([(0, 1, 0, 2)],)), ['0', '0', '0'])
check("variable forecast count",solve([(0,1,0,N),(1,1,N,0)]),["0","1/4","1/4"])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| regression 1 | ['0', '1/16', '1/4'] | ['0', '1/16', '1/4'] | Passed |
| regression 2 | ['0', '3/64', '15/64'] | ['0', '3/64', '15/64'] | Passed |
| regression 3 | None | None | Passed |
| regression 4 | ['0', '1/4', '1/4'] | ['0', '1/4', '1/4'] | Passed |
| regression 5 | ['16/225', '0', '6/25'] | ['16/225', '0', '6/25'] | Passed |
| regression 6 | ['23/75', '49/960', '15/64'] | ['23/75', '49/960', '15/64'] | Passed |
| regression 7 | ['0', '0', '0'] | ['0', '0', '0'] | Passed |
| variable forecast count | ['0', '1/4', '1/4'] | ['0', '1/4', '1/4'] | Passed |
SHA-256 / 03d2899f3cf3a0d2d5f43bff6823c5ec9b99d0dd4fcdb428051c0b6e4f23c702
Verification & scope
Small offline integer/rational inputs only; no performance, statistical inference, or production-library conformance claim. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:39:12.315698+00:00.
Case digest / cde845ec518f15b6f508596e026b65d77b8f3bc0334f998d4a6ab124209eaabd