FA-84261 / Sports scoring and tiebreakers / Open access
Trimmed marks averaged before multiplying by difficulty · case 01
Dive scores are a third of the expected value.
ROOT CAUSE
The kept marks are averaged instead of summed.
VERIFIED REPAIR
Sum the three kept marks and multiply by the degree of difficulty.
Unsuccessful approach: Averaging all marks and scaling by three ignores the trimming.
Case contract
Diving dive score. scores are judge marks as strings from 0 to 10 in half points; any other mark returns "invalid mark <s>". A 5-judge panel drops the single highest and lowest marks, a 7-judge panel the two highest and two lowest; any other panel size returns "invalid panel". The three remaining marks are summed and multiplied by the degree of difficulty dd (a decimal string); return the exact result with two decimals.
Why this case matters
Meet management systems compute dive scores from judge panels of different sizes.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v > 10 or (v * 2).denominator != 1:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 2
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(kept) / len(kept) * Fraction(dd)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['2.0', '7.5', '4.0', '4.5', '2.0'], '3.7'), '38.85'),
('variant scenario 1', (['7.5', '2.5', '8.0', '0.0', '4.5', '5.5', '6.0'], '3.4'), '54.40'),
('variant scenario 2', (['5.5', '7.0', '9.0', '2.5', '0.0', '1.5', '7.0'], '3.4'), '51.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['8.5', '4.5', '10.0', '10.0', '1.0'], '2.0'), '46.00'),
('variant scenario 1', (['2.5', '5.0', '5.5', '2.5', '1.5', '6.5'], '1.6'), 'invalid panel'),
('variant scenario 2', (['7.5', '8.5', '0.0', '7.0', '5.5', '6.0', '6.5'], '3.7'), '72.15')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['6.5', '2.0', '8.5', '0.0', '8.5'], '3.1'), '52.70'),
('variant scenario 1', (['1.5', '8.5', '2.0', '4.5', '8.5'], '3.1'), '46.50'),
('variant scenario 2', (['6.0', '0.5', '9.0', '1.0', '7.5', '3.5', '4.0'], '3.4'), '45.90')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['4.0', '4.0', '3.5', '5.5', '3.5'], '3.1'), '35.65'),
('variant scenario 1', (['5.0', '10.0', '6.0', '9.0', '0.0', '5.0'], '2.0'), 'invalid panel'),
('variant scenario 2', (['11.0', '3.0', '5.5', '5.5', '7.0'], '3.4'), 'invalid mark 11.0')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average',
(['0.0', '5.5', '3.0', '7.5', '7.5', '10.0', '5.5'], '3.7'),
'68.45'),
('variant scenario 1', (['6.5', '2.5', '7.0', '6.5', '7.5', '4.0'], '3.1'), 'invalid panel'),
('variant scenario 2', (['9.0', '7.5', '7.5', '0.5', '1.5'], '3.7'), '61.05')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 14.33 | 43.00 | Failed |
| control seven-judge panel | 25.32 | 75.95 | Failed |
| boundary perfect tens | 34.00 | 102.00 | Failed |
| boundary non-half mark | invalid mark 7.3 | invalid mark 7.3 | Passed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: sum versus average | 12.95 | 38.85 | Failed |
| variant scenario 1 | 18.13 | 54.40 | Failed |
| variant scenario 2 | 17.00 | 51.00 | Failed |
SHA-256 / dfb7f106877f7bce37af45ae58bb2fb32380f1758ae11d926a0118f4ae63e3a6
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v > 10 or (v * 2).denominator != 1:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 2
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(marks) * Fraction(dd) * 3 / len(marks)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['2.0', '7.5', '4.0', '4.5', '2.0'], '3.7'), '38.85'),
('variant scenario 1', (['7.5', '2.5', '8.0', '0.0', '4.5', '5.5', '6.0'], '3.4'), '54.40'),
('variant scenario 2', (['5.5', '7.0', '9.0', '2.5', '0.0', '1.5', '7.0'], '3.4'), '51.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['8.5', '4.5', '10.0', '10.0', '1.0'], '2.0'), '46.00'),
('variant scenario 1', (['2.5', '5.0', '5.5', '2.5', '1.5', '6.5'], '1.6'), 'invalid panel'),
('variant scenario 2', (['7.5', '8.5', '0.0', '7.0', '5.5', '6.0', '6.5'], '3.7'), '72.15')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['6.5', '2.0', '8.5', '0.0', '8.5'], '3.1'), '52.70'),
('variant scenario 1', (['1.5', '8.5', '2.0', '4.5', '8.5'], '3.1'), '46.50'),
('variant scenario 2', (['6.0', '0.5', '9.0', '1.0', '7.5', '3.5', '4.0'], '3.4'), '45.90')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['4.0', '4.0', '3.5', '5.5', '3.5'], '3.1'), '35.65'),
('variant scenario 1', (['5.0', '10.0', '6.0', '9.0', '0.0', '5.0'], '2.0'), 'invalid panel'),
('variant scenario 2', (['11.0', '3.0', '5.5', '5.5', '7.0'], '3.4'), 'invalid mark 11.0')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average',
(['0.0', '5.5', '3.0', '7.5', '7.5', '10.0', '5.5'], '3.7'),
'68.45'),
('variant scenario 1', (['6.5', '2.5', '7.0', '6.5', '7.5', '4.0'], '3.1'), 'invalid panel'),
('variant scenario 2', (['9.0', '7.5', '7.5', '0.5', '1.5'], '3.7'), '61.05')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 43.20 | 43.00 | Failed |
| control seven-judge panel | 73.74 | 75.95 | Failed |
| boundary perfect tens | 102.00 | 102.00 | Passed |
| boundary non-half mark | invalid mark 7.3 | invalid mark 7.3 | Passed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: sum versus average | 44.40 | 38.85 | Failed |
| variant scenario 1 | 49.54 | 54.40 | Failed |
| variant scenario 2 | 47.36 | 51.00 | Failed |
SHA-256 / 2315cbf166a9a7533062689eacfa2235a2c2ec3e8af80991c90ce013c6564734
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v > 10 or (v * 2).denominator != 1:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 2
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(kept) * Fraction(dd)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['2.0', '7.5', '4.0', '4.5', '2.0'], '3.7'), '38.85'),
('variant scenario 1', (['7.5', '2.5', '8.0', '0.0', '4.5', '5.5', '6.0'], '3.4'), '54.40'),
('variant scenario 2', (['5.5', '7.0', '9.0', '2.5', '0.0', '1.5', '7.0'], '3.4'), '51.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['8.5', '4.5', '10.0', '10.0', '1.0'], '2.0'), '46.00'),
('variant scenario 1', (['2.5', '5.0', '5.5', '2.5', '1.5', '6.5'], '1.6'), 'invalid panel'),
('variant scenario 2', (['7.5', '8.5', '0.0', '7.0', '5.5', '6.0', '6.5'], '3.7'), '72.15')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['6.5', '2.0', '8.5', '0.0', '8.5'], '3.1'), '52.70'),
('variant scenario 1', (['1.5', '8.5', '2.0', '4.5', '8.5'], '3.1'), '46.50'),
('variant scenario 2', (['6.0', '0.5', '9.0', '1.0', '7.5', '3.5', '4.0'], '3.4'), '45.90')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average', (['4.0', '4.0', '3.5', '5.5', '3.5'], '3.1'), '35.65'),
('variant scenario 1', (['5.0', '10.0', '6.0', '9.0', '0.0', '5.0'], '2.0'), 'invalid panel'),
('variant scenario 2', (['11.0', '3.0', '5.5', '5.5', '7.0'], '3.4'), 'invalid mark 11.0')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: sum versus average',
(['0.0', '5.5', '3.0', '7.5', '7.5', '10.0', '5.5'], '3.7'),
'68.45'),
('variant scenario 1', (['6.5', '2.5', '7.0', '6.5', '7.5', '4.0'], '3.1'), 'invalid panel'),
('variant scenario 2', (['9.0', '7.5', '7.5', '0.5', '1.5'], '3.7'), '61.05')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 43.00 | 43.00 | Passed |
| control seven-judge panel | 75.95 | 75.95 | Passed |
| boundary perfect tens | 102.00 | 102.00 | Passed |
| boundary non-half mark | invalid mark 7.3 | invalid mark 7.3 | Passed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: sum versus average | 38.85 | 38.85 | Passed |
| variant scenario 1 | 54.40 | 54.40 | Passed |
| variant scenario 2 | 51.00 | 51.00 | Passed |
SHA-256 / b27a5b0d774f12dc2534dfa639455aa18f32d2d4c9fb2a8691c7b9e33fbbb1ab
Verification & scope
Stipulated, bounded toy contract stated in the contract field; not a claim of conformance with any governing body rulebook or operator house rules. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:50:29.178769+00:00.
Case digest / c236dddcf00485bb4ffcd8644562305a9d134ea952a5f0199c15456422b6c88b