FA-84256 / Sports scoring and tiebreakers / Open access
Judge mark validation accepts off-grid or rejects a perfect ten · case 01
A 7.3 mark is accepted, or a legitimate 10.0 is rejected.
ROOT CAUSE
Validation only checks the 0-10 range, not the half-point grid.
THE FAILURE
Validation only checks the 0-10 range, not the half-point grid.
Unsuccessful approach: Adding the half-point check but tightening the range to below 10 rejects perfect marks.
Case contract
Diving dive score. scores are judge marks as strings from 0 to 10 in half points; any other mark returns "invalid mark <s>". A 5-judge panel drops the single highest and lowest marks, a 7-judge panel the two highest and two lowest; any other panel size returns "invalid panel". The three remaining marks are summed and multiplied by the degree of difficulty dd (a decimal string); return the exact result with two decimals.
Why this case matters
Meet management systems compute dive scores from judge panels of different sizes.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v > 10:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 2
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(kept) * Fraction(dd)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation', (['3.5', '10.0', '4.0', '3.5', '7.5'], '2.0'), '30.00'),
('regression: mark validation', (['7.3', '8.0', '3.5', '1.5', '0.5'], '2.0'), 'invalid mark 7.3'),
('variant scenario 1', (['1.5', '8.0', '3.0', '7.0', '7.5'], '3.1'), '54.25'),
('variant scenario 2',
(['-1.0', '2.5', '2.0', '3.5', '0.0', '7.5', '6.5'], '2.0'),
'invalid mark -1.0')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation',
(['1.5', '8.0', '9.0', '1.5', '7.0', '10.0', '0.0'], '1.6'),
'26.40'),
('regression: mark validation',
(['7.3', '1.5', '7.0', '1.5', '1.5', '0.0'], '2.8'),
'invalid mark 7.3'),
('variant scenario 1', (['8.5', '3.0', '8.5', '2.0', '4.0', '4.0'], '1.6'), 'invalid panel'),
('variant scenario 2', (['8.0', '3.0', '7.5', '3.0', '0.0', '7.5'], '2.0'), 'invalid panel')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation', (['9.0', '4.5', '10.0', '9.5', '4.5'], '3.4'), '78.20'),
('regression: mark validation',
(['7.3', '5.0', '4.5', '5.0', '3.5', '7.5', '0.5'], '1.6'),
'invalid mark 7.3'),
('variant scenario 1', (['0.0', '6.5', '8.0', '7.5', '5.0'], '2.8'), '53.20'),
('variant scenario 2', (['8.5', '4.5', '1.0', '6.5', '6.5'], '3.4'), '59.50')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation', (['0.0', '1.0', '2.0', '10.0', '8.0'], '1.6'), '17.60'),
('regression: mark validation',
(['7.3', '1.0', '4.0', '0.0', '7.0', '0.5', '2.0'], '2.0'),
'invalid mark 7.3'),
('variant scenario 1',
(['11.0', '2.5', '10.0', '6.0', '2.0', '6.0'], '2.0'),
'invalid mark 11.0'),
('variant scenario 2', (['6.0', '1.5', '6.0', '5.5', '10.0'], '1.6'), '28.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation',
(['10.0', '8.0', '8.0', '9.5', '0.5', '6.0'], '3.1'),
'invalid panel'),
('regression: mark validation',
(['7.3', '0.5', '6.5', '0.5', '3.0', '0.0', '9.5'], '2.0'),
'invalid mark 7.3'),
('variant scenario 1', (['8.0', '0.0', '4.5', '8.0', '5.5'], '2.0'), '36.00'),
('variant scenario 2', (['9.5', '8.0', '3.5', '7.5', '4.0', '9.0'], '3.1'), 'invalid panel')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 43.00 | 43.00 | Passed |
| control seven-judge panel | 75.95 | 75.95 | Passed |
| boundary perfect tens | 102.00 | 102.00 | Passed |
| boundary non-half mark | 42.00 | invalid mark 7.3 | Failed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: mark validation | 30.00 | 30.00 | Passed |
| regression: mark validation | 24.60 | invalid mark 7.3 | Failed |
| variant scenario 1 | 54.25 | 54.25 | Passed |
| variant scenario 2 | invalid mark -1.0 | invalid mark -1.0 | Passed |
SHA-256 / 393c8e00958cfdcefa084b23cd55219bc31a099c08721b8c03fb0addbf9ec6c2
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
marks = []
for s in scores:
v = Fraction(s)
if v < 0 or v >= 10 or (v * 2).denominator != 1:
return 'invalid mark ' + s
marks.append(v)
if len(marks) == 5:
drop = 1
elif len(marks) == 7:
drop = 2
else:
return 'invalid panel'
kept = sorted(marks)[drop:len(marks) - drop]
total = sum(kept) * Fraction(dd)
return '%.2f' % float(total)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
try:
return solve(*args)
except Exception as exc:
return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation', (['3.5', '10.0', '4.0', '3.5', '7.5'], '2.0'), '30.00'),
('regression: mark validation', (['7.3', '8.0', '3.5', '1.5', '0.5'], '2.0'), 'invalid mark 7.3'),
('variant scenario 1', (['1.5', '8.0', '3.0', '7.0', '7.5'], '3.1'), '54.25'),
('variant scenario 2',
(['-1.0', '2.5', '2.0', '3.5', '0.0', '7.5', '6.5'], '2.0'),
'invalid mark -1.0')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation',
(['1.5', '8.0', '9.0', '1.5', '7.0', '10.0', '0.0'], '1.6'),
'26.40'),
('regression: mark validation',
(['7.3', '1.5', '7.0', '1.5', '1.5', '0.0'], '2.8'),
'invalid mark 7.3'),
('variant scenario 1', (['8.5', '3.0', '8.5', '2.0', '4.0', '4.0'], '1.6'), 'invalid panel'),
('variant scenario 2', (['8.0', '3.0', '7.5', '3.0', '0.0', '7.5'], '2.0'), 'invalid panel')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation', (['9.0', '4.5', '10.0', '9.5', '4.5'], '3.4'), '78.20'),
('regression: mark validation',
(['7.3', '5.0', '4.5', '5.0', '3.5', '7.5', '0.5'], '1.6'),
'invalid mark 7.3'),
('variant scenario 1', (['0.0', '6.5', '8.0', '7.5', '5.0'], '2.8'), '53.20'),
('variant scenario 2', (['8.5', '4.5', '1.0', '6.5', '6.5'], '3.4'), '59.50')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation', (['0.0', '1.0', '2.0', '10.0', '8.0'], '1.6'), '17.60'),
('regression: mark validation',
(['7.3', '1.0', '4.0', '0.0', '7.0', '0.5', '2.0'], '2.0'),
'invalid mark 7.3'),
('variant scenario 1',
(['11.0', '2.5', '10.0', '6.0', '2.0', '6.0'], '2.0'),
'invalid mark 11.0'),
('variant scenario 2', (['6.0', '1.5', '6.0', '5.5', '10.0'], '1.6'), '28.00')],
[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
('control seven-judge panel',
(['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
'75.95'),
('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
('regression: mark validation',
(['10.0', '8.0', '8.0', '9.5', '0.5', '6.0'], '3.1'),
'invalid panel'),
('regression: mark validation',
(['7.3', '0.5', '6.5', '0.5', '3.0', '0.0', '9.5'], '2.0'),
'invalid mark 7.3'),
('variant scenario 1', (['8.0', '0.0', '4.5', '8.0', '5.5'], '2.0'), '36.00'),
('variant scenario 2', (['9.5', '8.0', '3.5', '7.5', '4.0', '9.0'], '3.1'), 'invalid panel')]]
for label, args, expected in cases[N - 1]:
check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| control five-judge panel | 43.00 | 43.00 | Passed |
| control seven-judge panel | 75.95 | 75.95 | Passed |
| boundary perfect tens | invalid mark 10.0 | 102.00 | Failed |
| boundary non-half mark | invalid mark 7.3 | invalid mark 7.3 | Passed |
| boundary six judges | invalid panel | invalid panel | Passed |
| regression: mark validation | invalid mark 10.0 | 30.00 | Failed |
| regression: mark validation | invalid mark 7.3 | invalid mark 7.3 | Passed |
| variant scenario 1 | 54.25 | 54.25 | Passed |
| variant scenario 2 | invalid mark -1.0 | invalid mark -1.0 | Passed |
SHA-256 / 6c4f6720c59f519780c51cbc9b270dfe6fd411bac48605fcc648a84884004035
HELD IN THE MEMBER ARCHIVE
The verified repair and its recorded checks are member-only.
This mechanism has 9 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.
Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.
Member access is invitation-based. Sign in with your invited account to inspect the repair.
Sign in to the archive ↗Verification & scope
Stipulated, bounded toy contract stated in the contract field; not a claim of conformance with any governing body rulebook or operator house rules. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:50:29.181793+00:00.
Case digest / 3e40edb2d4e44e4cbc62d2b61bb9a348f460bcd4cdb136b36a59b07e4300638b