FAILURE MAP
← Case archive

FA-84256 / Sports scoring and tiebreakers / Open access

Judge mark validation accepts off-grid or rejects a perfect ten · case 01

A 7.3 mark is accepted, or a legitimate 10.0 is rejected.

Verified by executionVariant 1 · 9 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

Validation only checks the 0-10 range, not the half-point grid.

THE FAILURE

Validation only checks the 0-10 range, not the half-point grid.

Unsuccessful approach: Adding the half-point check but tightening the range to below 10 rejects perfect marks.

Case contract

Diving dive score. scores are judge marks as strings from 0 to 10 in half points; any other mark returns "invalid mark <s>". A 5-judge panel drops the single highest and lowest marks, a 7-judge panel the two highest and two lowest; any other panel size returns "invalid panel". The three remaining marks are summed and multiplied by the degree of difficulty dd (a decimal string); return the exact result with two decimals.

Why this case matters

Meet management systems compute dive scores from judge panels of different sizes.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v > 10:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[drop:len(marks) - drop]
    total = sum(kept) * Fraction(dd)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation', (['3.5', '10.0', '4.0', '3.5', '7.5'], '2.0'), '30.00'),
  ('regression: mark validation', (['7.3', '8.0', '3.5', '1.5', '0.5'], '2.0'), 'invalid mark 7.3'),
  ('variant scenario 1', (['1.5', '8.0', '3.0', '7.0', '7.5'], '3.1'), '54.25'),
  ('variant scenario 2',
   (['-1.0', '2.5', '2.0', '3.5', '0.0', '7.5', '6.5'], '2.0'),
   'invalid mark -1.0')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation',
   (['1.5', '8.0', '9.0', '1.5', '7.0', '10.0', '0.0'], '1.6'),
   '26.40'),
  ('regression: mark validation',
   (['7.3', '1.5', '7.0', '1.5', '1.5', '0.0'], '2.8'),
   'invalid mark 7.3'),
  ('variant scenario 1', (['8.5', '3.0', '8.5', '2.0', '4.0', '4.0'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['8.0', '3.0', '7.5', '3.0', '0.0', '7.5'], '2.0'), 'invalid panel')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation', (['9.0', '4.5', '10.0', '9.5', '4.5'], '3.4'), '78.20'),
  ('regression: mark validation',
   (['7.3', '5.0', '4.5', '5.0', '3.5', '7.5', '0.5'], '1.6'),
   'invalid mark 7.3'),
  ('variant scenario 1', (['0.0', '6.5', '8.0', '7.5', '5.0'], '2.8'), '53.20'),
  ('variant scenario 2', (['8.5', '4.5', '1.0', '6.5', '6.5'], '3.4'), '59.50')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation', (['0.0', '1.0', '2.0', '10.0', '8.0'], '1.6'), '17.60'),
  ('regression: mark validation',
   (['7.3', '1.0', '4.0', '0.0', '7.0', '0.5', '2.0'], '2.0'),
   'invalid mark 7.3'),
  ('variant scenario 1',
   (['11.0', '2.5', '10.0', '6.0', '2.0', '6.0'], '2.0'),
   'invalid mark 11.0'),
  ('variant scenario 2', (['6.0', '1.5', '6.0', '5.5', '10.0'], '1.6'), '28.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation',
   (['10.0', '8.0', '8.0', '9.5', '0.5', '6.0'], '3.1'),
   'invalid panel'),
  ('regression: mark validation',
   (['7.3', '0.5', '6.5', '0.5', '3.0', '0.0', '9.5'], '2.0'),
   'invalid mark 7.3'),
  ('variant scenario 1', (['8.0', '0.0', '4.5', '8.0', '5.5'], '2.0'), '36.00'),
  ('variant scenario 2', (['9.5', '8.0', '3.5', '7.5', '4.0', '9.0'], '3.1'), 'invalid panel')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel43.0043.00Passed
control seven-judge panel75.9575.95Passed
boundary perfect tens102.00102.00Passed
boundary non-half mark42.00invalid mark 7.3Failed
boundary six judgesinvalid panelinvalid panelPassed
regression: mark validation30.0030.00Passed
regression: mark validation24.60invalid mark 7.3Failed
variant scenario 154.2554.25Passed
variant scenario 2invalid mark -1.0invalid mark -1.0Passed

SHA-256 / 393c8e00958cfdcefa084b23cd55219bc31a099c08721b8c03fb0addbf9ec6c2

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v >= 10 or (v * 2).denominator != 1:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[drop:len(marks) - drop]
    total = sum(kept) * Fraction(dd)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation', (['3.5', '10.0', '4.0', '3.5', '7.5'], '2.0'), '30.00'),
  ('regression: mark validation', (['7.3', '8.0', '3.5', '1.5', '0.5'], '2.0'), 'invalid mark 7.3'),
  ('variant scenario 1', (['1.5', '8.0', '3.0', '7.0', '7.5'], '3.1'), '54.25'),
  ('variant scenario 2',
   (['-1.0', '2.5', '2.0', '3.5', '0.0', '7.5', '6.5'], '2.0'),
   'invalid mark -1.0')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation',
   (['1.5', '8.0', '9.0', '1.5', '7.0', '10.0', '0.0'], '1.6'),
   '26.40'),
  ('regression: mark validation',
   (['7.3', '1.5', '7.0', '1.5', '1.5', '0.0'], '2.8'),
   'invalid mark 7.3'),
  ('variant scenario 1', (['8.5', '3.0', '8.5', '2.0', '4.0', '4.0'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['8.0', '3.0', '7.5', '3.0', '0.0', '7.5'], '2.0'), 'invalid panel')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation', (['9.0', '4.5', '10.0', '9.5', '4.5'], '3.4'), '78.20'),
  ('regression: mark validation',
   (['7.3', '5.0', '4.5', '5.0', '3.5', '7.5', '0.5'], '1.6'),
   'invalid mark 7.3'),
  ('variant scenario 1', (['0.0', '6.5', '8.0', '7.5', '5.0'], '2.8'), '53.20'),
  ('variant scenario 2', (['8.5', '4.5', '1.0', '6.5', '6.5'], '3.4'), '59.50')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation', (['0.0', '1.0', '2.0', '10.0', '8.0'], '1.6'), '17.60'),
  ('regression: mark validation',
   (['7.3', '1.0', '4.0', '0.0', '7.0', '0.5', '2.0'], '2.0'),
   'invalid mark 7.3'),
  ('variant scenario 1',
   (['11.0', '2.5', '10.0', '6.0', '2.0', '6.0'], '2.0'),
   'invalid mark 11.0'),
  ('variant scenario 2', (['6.0', '1.5', '6.0', '5.5', '10.0'], '1.6'), '28.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: mark validation',
   (['10.0', '8.0', '8.0', '9.5', '0.5', '6.0'], '3.1'),
   'invalid panel'),
  ('regression: mark validation',
   (['7.3', '0.5', '6.5', '0.5', '3.0', '0.0', '9.5'], '2.0'),
   'invalid mark 7.3'),
  ('variant scenario 1', (['8.0', '0.0', '4.5', '8.0', '5.5'], '2.0'), '36.00'),
  ('variant scenario 2', (['9.5', '8.0', '3.5', '7.5', '4.0', '9.0'], '3.1'), 'invalid panel')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel43.0043.00Passed
control seven-judge panel75.9575.95Passed
boundary perfect tensinvalid mark 10.0102.00Failed
boundary non-half markinvalid mark 7.3invalid mark 7.3Passed
boundary six judgesinvalid panelinvalid panelPassed
regression: mark validationinvalid mark 10.030.00Failed
regression: mark validationinvalid mark 7.3invalid mark 7.3Passed
variant scenario 154.2554.25Passed
variant scenario 2invalid mark -1.0invalid mark -1.0Passed

SHA-256 / 6c4f6720c59f519780c51cbc9b270dfe6fd411bac48605fcc648a84884004035

HELD IN THE MEMBER ARCHIVE

The verified repair and its recorded checks are member-only.

This mechanism has 9 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.

Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.

Member access is invitation-based. Sign in with your invited account to inspect the repair.

Sign in to the archive ↗

Verification & scope

Stipulated, bounded toy contract stated in the contract field; not a claim of conformance with any governing body rulebook or operator house rules. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:50:29.181793+00:00.

Case digest / 3e40edb2d4e44e4cbc62d2b61bb9a348f460bcd4cdb136b36a59b07e4300638b