FA-75606 / Text diff and three-way merge / Open access
Rename similarity pairing: one deleted file is renamed to several added files · case 01
A single deleted file appears as the source of multiple renames.
ROOT CAUSE
The greedy pass checks only whether the added file was used.
VERIFIED REPAIR
Skip a candidate when either file is already paired.
Unsuccessful approach: Checking only the deleted side lets one added file absorb several deleted files.
Case contract
For deleted and added files (path -> lines), a pair scores floor(100 * multiset-common-lines / max(len_a, len_b)); empty files never pair. Pairs scoring at least the threshold are taken greedily by descending score, ties by (deleted path, added path) ascending, each file used at most once. Return [deleted, added, score] triples sorted by deleted path.
Why this case matters
Rename and copy detection decides whether a delete plus an add is shown as a rename with a small diff.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import collections
N = 1
observations = []
def solve(deleted, added, threshold):
cands = []
for o, ol in deleted.items():
for a_, al in added.items():
if not ol or not al:
continue
common = sum((collections.Counter(ol) & collections.Counter(al)).values())
score = common * 100 // max(len(ol), len(al))
if score >= threshold:
cands.append((-score, o, a_))
cands.sort()
used_o, used_a, pairs = set(), set(), []
for neg, o, a_ in cands:
if a_ in used_a:
continue
used_o.add(o)
used_a.add(a_)
pairs.append([o, a_, -neg])
return sorted(pairs)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
cases = {
1: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 26], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 26], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
2: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 27], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 27], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
3: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 28], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 28], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
4: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 29], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 29], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
5: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 30], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 30], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
}[N]
for label, args, expected in cases:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| exact rename | [['old.c', 'new.c', 100]] | [['old.c', 'new.c', 100]] | Passed |
| edited rename above threshold | [['a.py', 'b.py', 60]] | [['a.py', 'b.py', 60]] | Passed |
| score exactly at the threshold | [['a.py', 'b.py', 50]] | [['a.py', 'b.py', 50]] | Passed |
| small file inside a big one | [] | [] | Passed |
| duplicate lines count once per copy | [] | [] | Passed |
| duplicate lines on both sides | [['d1', 'd2', 50]] | [['d1', 'd2', 50]] | Passed |
| score one below the threshold | [] | [] | Passed |
| best pairing wins | [['a', 'c', 80]] | [['a', 'c', 80]] | Passed |
| each file used once | [['a', 'c1', 100], ['a', 'c2', 100]] | [['a', 'c1', 100]] | Failed |
| each destination used once | [['a1', 'c', 100]] | [['a1', 'c', 100]] | Passed |
| lower score pairs after higher ones | [['x', 'p', 90], ['x', 'q', 100]] | [['x', 'q', 100], ['y', 'p', 50]] | Failed |
| empty files never pair | [] | [] | Passed |
SHA-256 / a9d4833a4d4828136dfe9a3a045efba2145be9991671e4986e632164872c45e1
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import collections
N = 1
observations = []
def solve(deleted, added, threshold):
cands = []
for o, ol in deleted.items():
for a_, al in added.items():
if not ol or not al:
continue
common = sum((collections.Counter(ol) & collections.Counter(al)).values())
score = common * 100 // max(len(ol), len(al))
if score >= threshold:
cands.append((-score, o, a_))
cands.sort()
used_o, used_a, pairs = set(), set(), []
for neg, o, a_ in cands:
if o in used_o:
continue
used_o.add(o)
used_a.add(a_)
pairs.append([o, a_, -neg])
return sorted(pairs)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
cases = {
1: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 26], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 26], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
2: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 27], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 27], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
3: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 28], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 28], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
4: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 29], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 29], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
5: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 30], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 30], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
}[N]
for label, args, expected in cases:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| exact rename | [['old.c', 'new.c', 100]] | [['old.c', 'new.c', 100]] | Passed |
| edited rename above threshold | [['a.py', 'b.py', 60]] | [['a.py', 'b.py', 60]] | Passed |
| score exactly at the threshold | [['a.py', 'b.py', 50]] | [['a.py', 'b.py', 50]] | Passed |
| small file inside a big one | [] | [] | Passed |
| duplicate lines count once per copy | [] | [] | Passed |
| duplicate lines on both sides | [['d1', 'd2', 50]] | [['d1', 'd2', 50]] | Passed |
| score one below the threshold | [] | [] | Passed |
| best pairing wins | [['a', 'c', 80], ['b', 'c', 70]] | [['a', 'c', 80]] | Failed |
| each file used once | [['a', 'c1', 100]] | [['a', 'c1', 100]] | Passed |
| each destination used once | [['a1', 'c', 100], ['a2', 'c', 100]] | [['a1', 'c', 100]] | Failed |
| lower score pairs after higher ones | [['x', 'q', 100], ['y', 'p', 50]] | [['x', 'q', 100], ['y', 'p', 50]] | Passed |
| empty files never pair | [] | [] | Passed |
SHA-256 / 6ba987a733ca4835e138654ba532445f7a3ea0f1aab0edaa3d708798af60f4a9
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
import collections
N = 1
observations = []
def solve(deleted, added, threshold):
cands = []
for o, ol in deleted.items():
for a_, al in added.items():
if not ol or not al:
continue
common = sum((collections.Counter(ol) & collections.Counter(al)).values())
score = common * 100 // max(len(ol), len(al))
if score >= threshold:
cands.append((-score, o, a_))
cands.sort()
used_o, used_a, pairs = set(), set(), []
for neg, o, a_ in cands:
if o in used_o or a_ in used_a:
continue
used_o.add(o)
used_a.add(a_)
pairs.append([o, a_, -neg])
return sorted(pairs)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
cases = {
1: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 26], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 26], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
2: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 27], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 27], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
3: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 28], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 28], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
4: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 29], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 29], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
5: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 30], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 30], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
}[N]
for label, args, expected in cases:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| exact rename | [['old.c', 'new.c', 100]] | [['old.c', 'new.c', 100]] | Passed |
| edited rename above threshold | [['a.py', 'b.py', 60]] | [['a.py', 'b.py', 60]] | Passed |
| score exactly at the threshold | [['a.py', 'b.py', 50]] | [['a.py', 'b.py', 50]] | Passed |
| small file inside a big one | [] | [] | Passed |
| duplicate lines count once per copy | [] | [] | Passed |
| duplicate lines on both sides | [['d1', 'd2', 50]] | [['d1', 'd2', 50]] | Passed |
| score one below the threshold | [] | [] | Passed |
| best pairing wins | [['a', 'c', 80]] | [['a', 'c', 80]] | Passed |
| each file used once | [['a', 'c1', 100]] | [['a', 'c1', 100]] | Passed |
| each destination used once | [['a1', 'c', 100]] | [['a1', 'c', 100]] | Passed |
| lower score pairs after higher ones | [['x', 'q', 100], ['y', 'p', 50]] | [['x', 'q', 100], ['y', 'p', 50]] | Passed |
| empty files never pair | [] | [] | Passed |
SHA-256 / 48109ecfda2e0e025e8d114dcf3870c4472f1d21e7cbafaee87394686b899092
Verification & scope
A deterministic, bounded teaching model of one diff, patch or merge rule with stipulated conventions; it is not a production diff or version-control implementation and makes no claim of conformance to any specific tool. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:49:08.279442+00:00.
Case digest / 46ffeb8faa12006c608198cea4f221a24252157d865f1b4f0204253b4432242c