FA-75611 / Text diff and three-way merge / Open access
Rename similarity pairing: candidates are paired in path order instead of by score · case 01
A weaker pairing claims a file before its best match is considered.
ROOT CAUSE
Candidates are sorted by path, not by descending score.
THE FAILURE
Candidates are sorted by path, not by descending score.
Unsuccessful approach: Reversing the sort takes the lowest scores first.
Case contract
For deleted and added files (path -> lines), a pair scores floor(100 * multiset-common-lines / max(len_a, len_b)); empty files never pair. Pairs scoring at least the threshold are taken greedily by descending score, ties by (deleted path, added path) ascending, each file used at most once. Return [deleted, added, score] triples sorted by deleted path.
Why this case matters
Rename and copy detection decides whether a delete plus an add is shown as a rename with a small diff.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import collections
N = 1
observations = []
def solve(deleted, added, threshold):
cands = []
for o, ol in deleted.items():
for a_, al in added.items():
if not ol or not al:
continue
common = sum((collections.Counter(ol) & collections.Counter(al)).values())
score = common * 100 // max(len(ol), len(al))
if score >= threshold:
cands.append((-score, o, a_))
cands.sort(key=lambda c: (c[1], c[2]))
used_o, used_a, pairs = set(), set(), []
for neg, o, a_ in cands:
if o in used_o or a_ in used_a:
continue
used_o.add(o)
used_a.add(a_)
pairs.append([o, a_, -neg])
return sorted(pairs)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
cases = {
1: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 26], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 26], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
2: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 27], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 27], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
3: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 28], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 28], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
4: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 29], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 29], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
5: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 30], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 30], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
}[N]
for label, args, expected in cases:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| exact rename | [['old.c', 'new.c', 100]] | [['old.c', 'new.c', 100]] | Passed |
| edited rename above threshold | [['a.py', 'b.py', 60]] | [['a.py', 'b.py', 60]] | Passed |
| score exactly at the threshold | [['a.py', 'b.py', 50]] | [['a.py', 'b.py', 50]] | Passed |
| small file inside a big one | [] | [] | Passed |
| duplicate lines count once per copy | [] | [] | Passed |
| duplicate lines on both sides | [['d1', 'd2', 50]] | [['d1', 'd2', 50]] | Passed |
| score one below the threshold | [] | [] | Passed |
| best pairing wins | [['a', 'c', 80]] | [['a', 'c', 80]] | Passed |
| each file used once | [['a', 'c1', 100]] | [['a', 'c1', 100]] | Passed |
| each destination used once | [['a1', 'c', 100]] | [['a1', 'c', 100]] | Passed |
| lower score pairs after higher ones | [['x', 'p', 90], ['y', 'q', 50]] | [['x', 'q', 100], ['y', 'p', 50]] | Failed |
| empty files never pair | [] | [] | Passed |
SHA-256 / b93acc8cd150cb7eff0f0ce91333c5ec2e76a6882452b71cdba0e7af122e18af
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import collections
N = 1
observations = []
def solve(deleted, added, threshold):
cands = []
for o, ol in deleted.items():
for a_, al in added.items():
if not ol or not al:
continue
common = sum((collections.Counter(ol) & collections.Counter(al)).values())
score = common * 100 // max(len(ol), len(al))
if score >= threshold:
cands.append((-score, o, a_))
cands.sort(reverse=True)
used_o, used_a, pairs = set(), set(), []
for neg, o, a_ in cands:
if o in used_o or a_ in used_a:
continue
used_o.add(o)
used_a.add(a_)
pairs.append([o, a_, -neg])
return sorted(pairs)
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
cases = {
1: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 26], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 26], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
2: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 27], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 27], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
3: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 28], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 28], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
4: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 29], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 29], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
5: [('exact rename', [{'old.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'new.c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['old.c', 'new.c', 100]]), ('edited rename above threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'x', 'y', 'z', 'w']}, 60], [['a.py', 'b.py', 60]]), ('score exactly at the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 50], [['a.py', 'b.py', 50]]), ('small file inside a big one', [{'small': ['l0', 'l1', 'l2']}, {'big': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], []), ('duplicate lines count once per copy', [{'dup': ['}', '}', '}', 'x']}, {'dup2': ['}', 'y', 'z', 'w']}, 30], []), ('duplicate lines on both sides', [{'d1': ['}', '}', '}', 'x']}, {'d2': ['}', '}', 'y', 'z']}, 30], [['d1', 'd2', 50]]), ('score one below the threshold', [{'a.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'b.py': ['l0', 'l1', 'l2', 'l3', 'l4', 'q', 'q', 'q', 'q', 'q']}, 51], []), ('best pairing wins', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'b': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'b1', 'b2', 'b3']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'c1', 'c2']}, 50], [['a', 'c', 80]]), ('each file used once', [{'a': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'c2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a', 'c1', 100]]), ('each destination used once', [{'a1': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'a2': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, {'c': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 50], [['a1', 'c', 100]]), ('lower score pairs after higher ones', [{'x': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9'], 'y': ['l0', 'l1', 'l2', 'l3', 'l4', 'y0', 'y1', 'y2', 'y3', 'y4']}, {'p': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'p'], 'q': ['l0', 'l1', 'l2', 'l3', 'l4', 'l5', 'l6', 'l7', 'l8', 'l9']}, 40], [['x', 'q', 100], ['y', 'p', 50]]), ('empty files never pair', [{'e': []}, {'f': [], 'g': ['z']}, 0], [])],
}[N]
for label, args, expected in cases:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| exact rename | [['old.c', 'new.c', 100]] | [['old.c', 'new.c', 100]] | Passed |
| edited rename above threshold | [['a.py', 'b.py', 60]] | [['a.py', 'b.py', 60]] | Passed |
| score exactly at the threshold | [['a.py', 'b.py', 50]] | [['a.py', 'b.py', 50]] | Passed |
| small file inside a big one | [] | [] | Passed |
| duplicate lines count once per copy | [] | [] | Passed |
| duplicate lines on both sides | [['d1', 'd2', 50]] | [['d1', 'd2', 50]] | Passed |
| score one below the threshold | [] | [] | Passed |
| best pairing wins | [['b', 'c', 70]] | [['a', 'c', 80]] | Failed |
| each file used once | [['a', 'c2', 100]] | [['a', 'c1', 100]] | Failed |
| each destination used once | [['a2', 'c', 100]] | [['a1', 'c', 100]] | Failed |
| lower score pairs after higher ones | [['x', 'p', 90], ['y', 'q', 50]] | [['x', 'q', 100], ['y', 'p', 50]] | Failed |
| empty files never pair | [] | [] | Passed |
SHA-256 / ab09d76cc14d0c931643dd86d19a1166b676bc4779c95944de48dfeae15d10c9
HELD IN THE MEMBER ARCHIVE
The verified repair and its recorded checks are member-only.
This mechanism has 12 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.
Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.
Member access is invitation-based. Sign in with your invited account to inspect the repair.
Sign in to the archive ↗Verification & scope
A deterministic, bounded teaching model of one diff, patch or merge rule with stipulated conventions; it is not a production diff or version-control implementation and makes no claim of conformance to any specific tool. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:49:08.344679+00:00.
Case digest / 4f989a1976ca4e7e2ef6f72c25bd567d0c45a0b106d0349dc388f88d1129a671