FAILURE MAP
← Case archive

FA-84261 / Sports scoring and tiebreakers / Open access

Trimmed marks averaged before multiplying by difficulty · case 01

Dive scores are a third of the expected value.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The kept marks are averaged instead of summed.

VERIFIED REPAIR

Sum the three kept marks and multiply by the degree of difficulty.

Unsuccessful approach: Averaging all marks and scaling by three ignores the trimming.

Case contract

Diving dive score. scores are judge marks as strings from 0 to 10 in half points; any other mark returns "invalid mark <s>". A 5-judge panel drops the single highest and lowest marks, a 7-judge panel the two highest and two lowest; any other panel size returns "invalid panel". The three remaining marks are summed and multiplied by the degree of difficulty dd (a decimal string); return the exact result with two decimals.

Why this case matters

Meet management systems compute dive scores from judge panels of different sizes.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v > 10 or (v * 2).denominator != 1:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[drop:len(marks) - drop]
    total = sum(kept) / len(kept) * Fraction(dd)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['2.0', '7.5', '4.0', '4.5', '2.0'], '3.7'), '38.85'),
  ('variant scenario 1', (['7.5', '2.5', '8.0', '0.0', '4.5', '5.5', '6.0'], '3.4'), '54.40'),
  ('variant scenario 2', (['5.5', '7.0', '9.0', '2.5', '0.0', '1.5', '7.0'], '3.4'), '51.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['8.5', '4.5', '10.0', '10.0', '1.0'], '2.0'), '46.00'),
  ('variant scenario 1', (['2.5', '5.0', '5.5', '2.5', '1.5', '6.5'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['7.5', '8.5', '0.0', '7.0', '5.5', '6.0', '6.5'], '3.7'), '72.15')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['6.5', '2.0', '8.5', '0.0', '8.5'], '3.1'), '52.70'),
  ('variant scenario 1', (['1.5', '8.5', '2.0', '4.5', '8.5'], '3.1'), '46.50'),
  ('variant scenario 2', (['6.0', '0.5', '9.0', '1.0', '7.5', '3.5', '4.0'], '3.4'), '45.90')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['4.0', '4.0', '3.5', '5.5', '3.5'], '3.1'), '35.65'),
  ('variant scenario 1', (['5.0', '10.0', '6.0', '9.0', '0.0', '5.0'], '2.0'), 'invalid panel'),
  ('variant scenario 2', (['11.0', '3.0', '5.5', '5.5', '7.0'], '3.4'), 'invalid mark 11.0')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average',
   (['0.0', '5.5', '3.0', '7.5', '7.5', '10.0', '5.5'], '3.7'),
   '68.45'),
  ('variant scenario 1', (['6.5', '2.5', '7.0', '6.5', '7.5', '4.0'], '3.1'), 'invalid panel'),
  ('variant scenario 2', (['9.0', '7.5', '7.5', '0.5', '1.5'], '3.7'), '61.05')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel14.3343.00Failed
control seven-judge panel25.3275.95Failed
boundary perfect tens34.00102.00Failed
boundary non-half markinvalid mark 7.3invalid mark 7.3Passed
boundary six judgesinvalid panelinvalid panelPassed
regression: sum versus average12.9538.85Failed
variant scenario 118.1354.40Failed
variant scenario 217.0051.00Failed

SHA-256 / dfb7f106877f7bce37af45ae58bb2fb32380f1758ae11d926a0118f4ae63e3a6

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v > 10 or (v * 2).denominator != 1:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[drop:len(marks) - drop]
    total = sum(marks) * Fraction(dd) * 3 / len(marks)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['2.0', '7.5', '4.0', '4.5', '2.0'], '3.7'), '38.85'),
  ('variant scenario 1', (['7.5', '2.5', '8.0', '0.0', '4.5', '5.5', '6.0'], '3.4'), '54.40'),
  ('variant scenario 2', (['5.5', '7.0', '9.0', '2.5', '0.0', '1.5', '7.0'], '3.4'), '51.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['8.5', '4.5', '10.0', '10.0', '1.0'], '2.0'), '46.00'),
  ('variant scenario 1', (['2.5', '5.0', '5.5', '2.5', '1.5', '6.5'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['7.5', '8.5', '0.0', '7.0', '5.5', '6.0', '6.5'], '3.7'), '72.15')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['6.5', '2.0', '8.5', '0.0', '8.5'], '3.1'), '52.70'),
  ('variant scenario 1', (['1.5', '8.5', '2.0', '4.5', '8.5'], '3.1'), '46.50'),
  ('variant scenario 2', (['6.0', '0.5', '9.0', '1.0', '7.5', '3.5', '4.0'], '3.4'), '45.90')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['4.0', '4.0', '3.5', '5.5', '3.5'], '3.1'), '35.65'),
  ('variant scenario 1', (['5.0', '10.0', '6.0', '9.0', '0.0', '5.0'], '2.0'), 'invalid panel'),
  ('variant scenario 2', (['11.0', '3.0', '5.5', '5.5', '7.0'], '3.4'), 'invalid mark 11.0')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average',
   (['0.0', '5.5', '3.0', '7.5', '7.5', '10.0', '5.5'], '3.7'),
   '68.45'),
  ('variant scenario 1', (['6.5', '2.5', '7.0', '6.5', '7.5', '4.0'], '3.1'), 'invalid panel'),
  ('variant scenario 2', (['9.0', '7.5', '7.5', '0.5', '1.5'], '3.7'), '61.05')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel43.2043.00Failed
control seven-judge panel73.7475.95Failed
boundary perfect tens102.00102.00Passed
boundary non-half markinvalid mark 7.3invalid mark 7.3Passed
boundary six judgesinvalid panelinvalid panelPassed
regression: sum versus average44.4038.85Failed
variant scenario 149.5454.40Failed
variant scenario 247.3651.00Failed

SHA-256 / 2315cbf166a9a7533062689eacfa2235a2c2ec3e8af80991c90ce013c6564734

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v > 10 or (v * 2).denominator != 1:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[drop:len(marks) - drop]
    total = sum(kept) * Fraction(dd)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['2.0', '7.5', '4.0', '4.5', '2.0'], '3.7'), '38.85'),
  ('variant scenario 1', (['7.5', '2.5', '8.0', '0.0', '4.5', '5.5', '6.0'], '3.4'), '54.40'),
  ('variant scenario 2', (['5.5', '7.0', '9.0', '2.5', '0.0', '1.5', '7.0'], '3.4'), '51.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['8.5', '4.5', '10.0', '10.0', '1.0'], '2.0'), '46.00'),
  ('variant scenario 1', (['2.5', '5.0', '5.5', '2.5', '1.5', '6.5'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['7.5', '8.5', '0.0', '7.0', '5.5', '6.0', '6.5'], '3.7'), '72.15')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['6.5', '2.0', '8.5', '0.0', '8.5'], '3.1'), '52.70'),
  ('variant scenario 1', (['1.5', '8.5', '2.0', '4.5', '8.5'], '3.1'), '46.50'),
  ('variant scenario 2', (['6.0', '0.5', '9.0', '1.0', '7.5', '3.5', '4.0'], '3.4'), '45.90')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average', (['4.0', '4.0', '3.5', '5.5', '3.5'], '3.1'), '35.65'),
  ('variant scenario 1', (['5.0', '10.0', '6.0', '9.0', '0.0', '5.0'], '2.0'), 'invalid panel'),
  ('variant scenario 2', (['11.0', '3.0', '5.5', '5.5', '7.0'], '3.4'), 'invalid mark 11.0')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: sum versus average',
   (['0.0', '5.5', '3.0', '7.5', '7.5', '10.0', '5.5'], '3.7'),
   '68.45'),
  ('variant scenario 1', (['6.5', '2.5', '7.0', '6.5', '7.5', '4.0'], '3.1'), 'invalid panel'),
  ('variant scenario 2', (['9.0', '7.5', '7.5', '0.5', '1.5'], '3.7'), '61.05')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel43.0043.00Passed
control seven-judge panel75.9575.95Passed
boundary perfect tens102.00102.00Passed
boundary non-half markinvalid mark 7.3invalid mark 7.3Passed
boundary six judgesinvalid panelinvalid panelPassed
regression: sum versus average38.8538.85Passed
variant scenario 154.4054.40Passed
variant scenario 251.0051.00Passed

SHA-256 / b27a5b0d774f12dc2534dfa639455aa18f32d2d4c9fb2a8691c7b9e33fbbb1ab

Verification & scope

Stipulated, bounded toy contract stated in the contract field; not a claim of conformance with any governing body rulebook or operator house rules. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:50:29.178769+00:00.

Case digest / c236dddcf00485bb4ffcd8644562305a9d134ea952a5f0199c15456422b6c88b