FA-11931 / Scientific pipeline provenance / Open access
Technical repeats inflate biological replicate counts · case 01
Technical repeats inflate biological replicate counts.
ROOT CAUSE
Measurement records are counted as independent specimens.
VERIFIED REPAIR
Count distinct biological donor IDs within each condition.
Unsuccessful approach: Deduplicating aliquot IDs still counts multiple aliquots from one donor.
Case contract
Rows are condition, donor, aliquot tuples; return a dictionary of distinct donor counts per condition, including donors shared across conditions independently.
Why this case matters
A deterministic offline model of scientific workflow bookkeeping; the fixtures test provenance contracts without modeling instruments or biological inference.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(rows):
out={}
for condition,donor,aliquot in rows: out[condition]=out.get(condition,0)+1
return out
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
d='donor'+str(N)
check('two aliquots one donor', solve([('c',d,'a'),('c',d,'b')]), {'c':1})
check('repeated assay', solve([('c',d,'a'),('c',d,'a')]), {'c':1})
check('two donors', solve([('c',d,'a'),('c','other','b')]), {'c':2})
check('donor spans conditions', solve([('c',d,'a'),('t',d,'b')]), {'c':1,'t':1})
check('empty experiment', solve([]), {})
check('aliquot names local to donor', solve([('c',d,'a'),('c','other','a')]), {'c':2})
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| two aliquots one donor | {'c': 2} | {'c': 1} | Failed |
| repeated assay | {'c': 2} | {'c': 1} | Failed |
| two donors | {'c': 2} | {'c': 2} | Passed |
| donor spans conditions | {'c': 1, 't': 1} | {'c': 1, 't': 1} | Passed |
| empty experiment | {} | {} | Passed |
| aliquot names local to donor | {'c': 2} | {'c': 2} | Passed |
SHA-256 / 64892847795abeb04cd92adadbb07924f1b0b45cf9a7da00fb67db1955f5fe0d
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(rows):
out={}
for condition,donor,aliquot in rows: out.setdefault(condition,set()).add(aliquot)
return {k:len(v) for k,v in out.items()}
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
d='donor'+str(N)
check('two aliquots one donor', solve([('c',d,'a'),('c',d,'b')]), {'c':1})
check('repeated assay', solve([('c',d,'a'),('c',d,'a')]), {'c':1})
check('two donors', solve([('c',d,'a'),('c','other','b')]), {'c':2})
check('donor spans conditions', solve([('c',d,'a'),('t',d,'b')]), {'c':1,'t':1})
check('empty experiment', solve([]), {})
check('aliquot names local to donor', solve([('c',d,'a'),('c','other','a')]), {'c':2})
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| two aliquots one donor | {'c': 2} | {'c': 1} | Failed |
| repeated assay | {'c': 1} | {'c': 1} | Passed |
| two donors | {'c': 2} | {'c': 2} | Passed |
| donor spans conditions | {'c': 1, 't': 1} | {'c': 1, 't': 1} | Passed |
| empty experiment | {} | {} | Passed |
| aliquot names local to donor | {'c': 1} | {'c': 2} | Failed |
SHA-256 / 874d964738d49ccfc6795649d320eed294012e109cdf64db4b05c8cf02f48796
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(rows):
out={}
for condition,donor,aliquot in rows: out.setdefault(condition,set()).add(donor)
return {k:len(v) for k,v in out.items()}
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
d='donor'+str(N)
check('two aliquots one donor', solve([('c',d,'a'),('c',d,'b')]), {'c':1})
check('repeated assay', solve([('c',d,'a'),('c',d,'a')]), {'c':1})
check('two donors', solve([('c',d,'a'),('c','other','b')]), {'c':2})
check('donor spans conditions', solve([('c',d,'a'),('t',d,'b')]), {'c':1,'t':1})
check('empty experiment', solve([]), {})
check('aliquot names local to donor', solve([('c',d,'a'),('c','other','a')]), {'c':2})
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| two aliquots one donor | {'c': 1} | {'c': 1} | Passed |
| repeated assay | {'c': 1} | {'c': 1} | Passed |
| two donors | {'c': 2} | {'c': 2} | Passed |
| donor spans conditions | {'c': 1, 't': 1} | {'c': 1, 't': 1} | Passed |
| empty experiment | {} | {} | Passed |
| aliquot names local to donor | {'c': 2} | {'c': 2} | Passed |
SHA-256 / 02af0e3188f77fb2a857b5260572f37e6ba3449aa433ec672a28c7aa1088c32a
Verification & scope
In-memory symbolic records only; no instrument, assay, or production workflow validation. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:38:52.329402+00:00.
Case digest / 104853f18afb1a89ba9e248c3adcd9f57cdb3bd734cb33fb4e865018878d4562