FAILURE MAP
← Case archive

FA-84251 / Sports scoring and tiebreakers / Open access

Only the low marks are trimmed · case 01

A generous outlier judge inflates the dive score.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The slice removes marks from the low end only.

VERIFIED REPAIR

Slice from drop to len - drop.

Unsuccessful approach: Trimming only the high end lets a harsh outlier deflate the score.

Case contract

Diving dive score. scores are judge marks as strings from 0 to 10 in half points; any other mark returns "invalid mark <s>". A 5-judge panel drops the single highest and lowest marks, a 7-judge panel the two highest and two lowest; any other panel size returns "invalid panel". The three remaining marks are summed and multiplied by the degree of difficulty dd (a decimal string); return the exact result with two decimals.

Why this case matters

Meet management systems compute dive scores from judge panels of different sizes.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v > 10 or (v * 2).denominator != 1:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[drop:]
    total = sum(kept) * Fraction(dd)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['8.0', '2.0', '4.5', '3.5', '0.5', '2.5', '6.5'], '1.6'),
   '16.80'),
  ('variant scenario 1', (['4.0', '4.5', '0.0', '9.0', '7.0', '1.5'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['8.5', '8.5', '1.5', '1.0', '4.0'], '1.6'), '22.40')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['10.0', '8.5', '3.5', '2.0', '3.0', '2.5', '8.0'], '2.8'),
   '40.60'),
  ('variant scenario 1', (['1.5', '6.5', '10.0', '8.5', '0.0'], '3.1'), '51.15'),
  ('variant scenario 2', (['7.3', '3.5', '5.0', '6.5', '7.5', '2.0'], '3.1'), 'invalid mark 7.3')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['10.0', '2.0', '5.5', '2.0', '9.5', '5.5', '9.5'], '3.1'),
   '63.55'),
  ('variant scenario 1', (['-1.0', '5.5', '3.5', '5.0', '8.5', '9.5'], '1.6'), 'invalid mark -1.0'),
  ('variant scenario 2', (['1.0', '1.5', '6.5', '8.0', '7.0', '7.0', '4.0'], '1.6'), '28.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends', (['4.0', '0.5', '2.5', '1.5', '0.0'], '3.1'), '13.95'),
  ('regression: trim both ends', (['8.5', '9.5', '3.0', '3.5', '8.5'], '1.6'), '32.80'),
  ('variant scenario 1', (['1.0', '7.5', '1.5', '1.0', '0.0'], '3.1'), '10.85'),
  ('variant scenario 2', (['8.0', '2.0', '3.0', '7.5', '7.0', '5.0'], '3.1'), 'invalid panel')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends', (['8.5', '6.5', '9.5', '5.5', '7.5'], '2.0'), '45.00'),
  ('variant scenario 1', (['2.5', '3.0', '8.5', '7.5', '4.5', '2.0'], '3.1'), 'invalid panel'),
  ('variant scenario 2', (['10.5', '4.5', '0.5', '3.5', '9.5'], '3.1'), 'invalid mark 10.5')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel59.0043.00Failed
control seven-judge panel130.2075.95Failed
boundary perfect tens136.00102.00Failed
boundary non-half markinvalid mark 7.3invalid mark 7.3Passed
boundary six judgesinvalid panelinvalid panelPassed
regression: trim both ends40.0016.80Failed
variant scenario 1invalid panelinvalid panelPassed
variant scenario 236.0022.40Failed

SHA-256 / 1b340dabfbcc320d937f95d76867a3264bdf05c4a75371cecfaeb7826fd60a92

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v > 10 or (v * 2).denominator != 1:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[:len(marks) - drop]
    total = sum(kept) * Fraction(dd)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['8.0', '2.0', '4.5', '3.5', '0.5', '2.5', '6.5'], '1.6'),
   '16.80'),
  ('variant scenario 1', (['4.0', '4.5', '0.0', '9.0', '7.0', '1.5'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['8.5', '8.5', '1.5', '1.0', '4.0'], '1.6'), '22.40')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['10.0', '8.5', '3.5', '2.0', '3.0', '2.5', '8.0'], '2.8'),
   '40.60'),
  ('variant scenario 1', (['1.5', '6.5', '10.0', '8.5', '0.0'], '3.1'), '51.15'),
  ('variant scenario 2', (['7.3', '3.5', '5.0', '6.5', '7.5', '2.0'], '3.1'), 'invalid mark 7.3')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['10.0', '2.0', '5.5', '2.0', '9.5', '5.5', '9.5'], '3.1'),
   '63.55'),
  ('variant scenario 1', (['-1.0', '5.5', '3.5', '5.0', '8.5', '9.5'], '1.6'), 'invalid mark -1.0'),
  ('variant scenario 2', (['1.0', '1.5', '6.5', '8.0', '7.0', '7.0', '4.0'], '1.6'), '28.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends', (['4.0', '0.5', '2.5', '1.5', '0.0'], '3.1'), '13.95'),
  ('regression: trim both ends', (['8.5', '9.5', '3.0', '3.5', '8.5'], '1.6'), '32.80'),
  ('variant scenario 1', (['1.0', '7.5', '1.5', '1.0', '0.0'], '3.1'), '10.85'),
  ('variant scenario 2', (['8.0', '2.0', '3.0', '7.5', '7.0', '5.0'], '3.1'), 'invalid panel')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends', (['8.5', '6.5', '9.5', '5.5', '7.5'], '2.0'), '45.00'),
  ('variant scenario 1', (['2.5', '3.0', '8.5', '7.5', '4.5', '2.0'], '3.1'), 'invalid panel'),
  ('variant scenario 2', (['10.5', '4.5', '0.5', '3.5', '9.5'], '3.1'), 'invalid mark 10.5')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel56.0043.00Failed
control seven-judge panel117.8075.95Failed
boundary perfect tens136.00102.00Failed
boundary non-half markinvalid mark 7.3invalid mark 7.3Passed
boundary six judgesinvalid panelinvalid panelPassed
regression: trim both ends20.8016.80Failed
variant scenario 1invalid panelinvalid panelPassed
variant scenario 224.0022.40Failed

SHA-256 / 6e5d27408d7c87994a8ad0e15245a1170b31e9988fd9b7504498f6d42fdee877

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
from fractions import Fraction
N = 1
observations = []
def solve(scores, dd):
    marks = []
    for s in scores:
        v = Fraction(s)
        if v < 0 or v > 10 or (v * 2).denominator != 1:
            return 'invalid mark ' + s
        marks.append(v)
    if len(marks) == 5:
        drop = 1
    elif len(marks) == 7:
        drop = 2
    else:
        return 'invalid panel'
    kept = sorted(marks)[drop:len(marks) - drop]
    total = sum(kept) * Fraction(dd)
    return '%.2f' % float(total)
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
def run(args):
    try:
        return solve(*args)
    except Exception as exc:
        return 'raised ' + type(exc).__name__
cases = [[('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['8.0', '2.0', '4.5', '3.5', '0.5', '2.5', '6.5'], '1.6'),
   '16.80'),
  ('variant scenario 1', (['4.0', '4.5', '0.0', '9.0', '7.0', '1.5'], '1.6'), 'invalid panel'),
  ('variant scenario 2', (['8.5', '8.5', '1.5', '1.0', '4.0'], '1.6'), '22.40')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['10.0', '8.5', '3.5', '2.0', '3.0', '2.5', '8.0'], '2.8'),
   '40.60'),
  ('variant scenario 1', (['1.5', '6.5', '10.0', '8.5', '0.0'], '3.1'), '51.15'),
  ('variant scenario 2', (['7.3', '3.5', '5.0', '6.5', '7.5', '2.0'], '3.1'), 'invalid mark 7.3')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends',
   (['10.0', '2.0', '5.5', '2.0', '9.5', '5.5', '9.5'], '3.1'),
   '63.55'),
  ('variant scenario 1', (['-1.0', '5.5', '3.5', '5.0', '8.5', '9.5'], '1.6'), 'invalid mark -1.0'),
  ('variant scenario 2', (['1.0', '1.5', '6.5', '8.0', '7.0', '7.0', '4.0'], '1.6'), '28.00')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends', (['4.0', '0.5', '2.5', '1.5', '0.0'], '3.1'), '13.95'),
  ('regression: trim both ends', (['8.5', '9.5', '3.0', '3.5', '8.5'], '1.6'), '32.80'),
  ('variant scenario 1', (['1.0', '7.5', '1.5', '1.0', '0.0'], '3.1'), '10.85'),
  ('variant scenario 2', (['8.0', '2.0', '3.0', '7.5', '7.0', '5.0'], '3.1'), 'invalid panel')],
 [('control five-judge panel', (['7.0', '7.5', '6.5', '8.0', '7.0'], '2.0'), '43.00'),
  ('control seven-judge panel',
   (['8.0', '8.5', '9.0', '7.5', '8.0', '8.5', '6.0'], '3.1'),
   '75.95'),
  ('boundary perfect tens', (['10.0', '10.0', '10.0', '10.0', '10.0'], '3.4'), '102.00'),
  ('boundary non-half mark', (['7.3', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid mark 7.3'),
  ('boundary six judges', (['7.0', '7.0', '7.0', '7.0', '7.0', '7.0'], '2.0'), 'invalid panel'),
  ('regression: trim both ends', (['8.5', '6.5', '9.5', '5.5', '7.5'], '2.0'), '45.00'),
  ('variant scenario 1', (['2.5', '3.0', '8.5', '7.5', '4.5', '2.0'], '3.1'), 'invalid panel'),
  ('variant scenario 2', (['10.5', '4.5', '0.5', '3.5', '9.5'], '3.1'), 'invalid mark 10.5')]]
for label, args, expected in cases[N - 1]:
    check(label, run(args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
control five-judge panel43.0043.00Passed
control seven-judge panel75.9575.95Passed
boundary perfect tens102.00102.00Passed
boundary non-half markinvalid mark 7.3invalid mark 7.3Passed
boundary six judgesinvalid panelinvalid panelPassed
regression: trim both ends16.8016.80Passed
variant scenario 1invalid panelinvalid panelPassed
variant scenario 222.4022.40Passed

SHA-256 / b93259cecfe1534eea3d1c115cde25af921009a68e71eebf89e9700e30d376ec

Verification & scope

Stipulated, bounded toy contract stated in the contract field; not a claim of conformance with any governing body rulebook or operator house rules. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:50:29.144485+00:00.

Case digest / f6d489313785f498a49dc9a16d1ad24bc66b48407936c67cedb8b0b45a2f8b1c