FA-84246 / Sports scoring and tiebreakers / Open access
Seven-judge panel trimmed like a five-judge panel · case 01
Seven-judge dive scores sum five marks and come out far too high.
ROOT CAUSE
The drop count for seven judges is 1.
VERIFIED REPAIR
Drop two high and two low marks for seven judges.
Unsuccessful approach: Dropping three from each end keeps only the median mark.
Case contract
Diving dive score. scores are judge marks as strings from 0 to 10 in half points; any other mark returns "invalid mark <s>". A 5-judge panel drops the single highest and lowest marks, a 7-judge panel the two highest and two lowest; any other panel size returns "invalid panel". The three remaining marks are summed and multiplied by the degree of difficulty dd (a decimal string); return the exact result with two decimals.
Why this case matters
Meet management systems compute dive scores from judge panels of different sizes.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v > 10 or (v * 2).denominator != 1:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 1
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(kept) * Fraction(dd)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['5.5', '1.5', '4.0', '1.5', '7.5', '0.5', '9.0'], '3.7'),
'40.70'),
('variant scenario 1', (['8.5', '4.5', '2.5', '3.5', '5.5'], '3.4'), '45.90'),
('variant scenario 2', (['4.0', '3.0', '8.0', '8.5', '1.0'], '2.0'), '30.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['3.5', '9.0', '5.5', '2.5', '3.0', '5.5', '4.0'], '3.4'),
'44.20'),
('variant scenario 1',
(['7.3', '4.5', '4.5', '6.0', '0.0', '5.0', '0.5'], '3.1'),
'invalid mark 7.3'),
('variant scenario 2', (['10.0', '8.0', '3.0', '3.0', '9.0'], '3.7'), '74.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['2.5', '6.5', '9.0', '3.5', '6.0', '0.5', '9.0'], '3.7'),
'59.20'),
('variant scenario 1', (['1.5', '6.5', '10.0', '2.5', '8.0', '8.5'], '3.1'), 'invalid panel'),
('variant scenario 2', (['8.0', '0.0', '6.5', '9.0', '5.0', '4.0', '6.0'], '3.7'), '64.75')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['7.5', '7.0', '10.0', '0.0', '1.5', '2.5', '9.0'], '1.6'),
'27.20'),
('variant scenario 1', (['10.0', '2.5', '7.0', '1.0', '9.5'], '1.6'), '30.40'),
('variant scenario 2', (['1.5', '0.5', '7.5', '8.0', '7.0', '1.0'], '3.7'), 'invalid panel')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['1.5', '10.0', '6.0', '4.0', '9.0', '4.5', '0.0'], '3.1'),
'44.95'),
('variant scenario 1', (['4.5', '5.5', '1.5', '10.0', '9.5', '9.5'], '2.8'), 'invalid panel'),
('variant scenario 2', (['8.5', '5.5', '4.0', '0.5', '10.0'], '3.4'), '61.20')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 43.00 | 43.00 | Passed |
| control seven-judge panel | 125.55 | 75.95 | Failed |
| boundary perfect tens | 102.00 | 102.00 | Passed |
| boundary non-half mark | invalid mark 7.3 | invalid mark 7.3 | Passed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: seven judge trim | 74.00 | 40.70 | Failed |
| variant scenario 1 | 45.90 | 45.90 | Passed |
| variant scenario 2 | 30.00 | 30.00 | Passed |
SHA-256 / 68ab8c63011d35ec09c52729e6ac8370195978c03b346fc5098f05d3de0fc7af
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v > 10 or (v * 2).denominator != 1:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 3
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(kept) * Fraction(dd)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['5.5', '1.5', '4.0', '1.5', '7.5', '0.5', '9.0'], '3.7'),
'40.70'),
('variant scenario 1', (['8.5', '4.5', '2.5', '3.5', '5.5'], '3.4'), '45.90'),
('variant scenario 2', (['4.0', '3.0', '8.0', '8.5', '1.0'], '2.0'), '30.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['3.5', '9.0', '5.5', '2.5', '3.0', '5.5', '4.0'], '3.4'),
'44.20'),
('variant scenario 1',
(['7.3', '4.5', '4.5', '6.0', '0.0', '5.0', '0.5'], '3.1'),
'invalid mark 7.3'),
('variant scenario 2', (['10.0', '8.0', '3.0', '3.0', '9.0'], '3.7'), '74.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['2.5', '6.5', '9.0', '3.5', '6.0', '0.5', '9.0'], '3.7'),
'59.20'),
('variant scenario 1', (['1.5', '6.5', '10.0', '2.5', '8.0', '8.5'], '3.1'), 'invalid panel'),
('variant scenario 2', (['8.0', '0.0', '6.5', '9.0', '5.0', '4.0', '6.0'], '3.7'), '64.75')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['7.5', '7.0', '10.0', '0.0', '1.5', '2.5', '9.0'], '1.6'),
'27.20'),
('variant scenario 1', (['10.0', '2.5', '7.0', '1.0', '9.5'], '1.6'), '30.40'),
('variant scenario 2', (['1.5', '0.5', '7.5', '8.0', '7.0', '1.0'], '3.7'), 'invalid panel')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['1.5', '10.0', '6.0', '4.0', '9.0', '4.5', '0.0'], '3.1'),
'44.95'),
('variant scenario 1', (['4.5', '5.5', '1.5', '10.0', '9.5', '9.5'], '2.8'), 'invalid panel'),
('variant scenario 2', (['8.5', '5.5', '4.0', '0.5', '10.0'], '3.4'), '61.20')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 43.00 | 43.00 | Passed |
| control seven-judge panel | 24.80 | 75.95 | Failed |
| boundary perfect tens | 102.00 | 102.00 | Passed |
| boundary non-half mark | invalid mark 7.3 | invalid mark 7.3 | Passed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: seven judge trim | 14.80 | 40.70 | Failed |
| variant scenario 1 | 45.90 | 45.90 | Passed |
| variant scenario 2 | 30.00 | 30.00 | Passed |
SHA-256 / 2253001aa712b6d40883bbb4772dc59a710f3a44807613f37bb70788c127e74c
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v > 10 or (v * 2).denominator != 1:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 2
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(kept) * Fraction(dd)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['5.5', '1.5', '4.0', '1.5', '7.5', '0.5', '9.0'], '3.7'),
'40.70'),
('variant scenario 1', (['8.5', '4.5', '2.5', '3.5', '5.5'], '3.4'), '45.90'),
('variant scenario 2', (['4.0', '3.0', '8.0', '8.5', '1.0'], '2.0'), '30.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['3.5', '9.0', '5.5', '2.5', '3.0', '5.5', '4.0'], '3.4'),
'44.20'),
('variant scenario 1',
(['7.3', '4.5', '4.5', '6.0', '0.0', '5.0', '0.5'], '3.1'),
'invalid mark 7.3'),
('variant scenario 2', (['10.0', '8.0', '3.0', '3.0', '9.0'], '3.7'), '74.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['2.5', '6.5', '9.0', '3.5', '6.0', '0.5', '9.0'], '3.7'),
'59.20'),
('variant scenario 1', (['1.5', '6.5', '10.0', '2.5', '8.0', '8.5'], '3.1'), 'invalid panel'),
('variant scenario 2', (['8.0', '0.0', '6.5', '9.0', '5.0', '4.0', '6.0'], '3.7'), '64.75')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['7.5', '7.0', '10.0', '0.0', '1.5', '2.5', '9.0'], '1.6'),
'27.20'),
('variant scenario 1', (['10.0', '2.5', '7.0', '1.0', '9.5'], '1.6'), '30.40'),
('variant scenario 2', (['1.5', '0.5', '7.5', '8.0', '7.0', '1.0'], '3.7'), 'invalid panel')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: seven judge trim',
(['1.5', '10.0', '6.0', '4.0', '9.0', '4.5', '0.0'], '3.1'),
'44.95'),
('variant scenario 1', (['4.5', '5.5', '1.5', '10.0', '9.5', '9.5'], '2.8'), 'invalid panel'),
('variant scenario 2', (['8.5', '5.5', '4.0', '0.5', '10.0'], '3.4'), '61.20')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 43.00 | 43.00 | Passed |
| control seven-judge panel | 75.95 | 75.95 | Passed |
| boundary perfect tens | 102.00 | 102.00 | Passed |
| boundary non-half mark | invalid mark 7.3 | invalid mark 7.3 | Passed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: seven judge trim | 40.70 | 40.70 | Passed |
| variant scenario 1 | 45.90 | 45.90 | Passed |
| variant scenario 2 | 30.00 | 30.00 | Passed |
SHA-256 / 6e3b06cd020371d144704e782842d94c8467c2bdaddcaaeb20a51078d107f2e9
Verification & scope
Stipulated, bounded toy contract stated in the contract field; not a claim of conformance with any governing body rulebook or operator house rules. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:50:29.144522+00:00.
Case digest / b238622b9126f4cae2b482eb1c9b9a7bcac86700ce45887c50cc8be57a4f6040