FA-47361 / Delimited text / Open access
Zipping columns silently discards missing or surplus values · case 01
A structured table violates the declared record or column contract.
ROOT CAUSE
Only one direction of schema width mismatch is checked.
THE FAILURE
Only one direction of schema width mismatch is checked.
Unsuccessful approach: Reversing the inequality merely exchanges truncation for missing values.
Case contract
Bind a comma-split header and data row. Header keys are ASCII case insensitive and stripped. Reject empty or duplicate canonical names, reserved _extra, and unequal width. Preserve data cell whitespace.
Why this case matters
Delimited interchange needs explicit framing, schema and field semantics at ingestion and emission boundaries.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
def _vary(value):
if value == '@END': return 3 + 4*N
if isinstance(value, str): return value.replace('@', 'cell' * N)
if isinstance(value, list): return [_vary(x) for x in value]
if isinstance(value, dict): return {_vary(k): _vary(v) for k,v in value.items()}
return value
N = 1
observations = []
def solve(data):
headers = [h.strip().lower() for h in data[0].split(',')]
values = data[1].split(',')
if any(not h for h in headers): return None
if len(set(headers)) != len(headers): return None
if '_extra' in headers: return None
if len(values) < len(headers): return None
return dict(zip(headers, values))
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('canonical names', solve(_vary([' A ,B', '@, y '])), _vary({'a': '@', 'b': ' y '}))
check('duplicate canonical', solve(_vary(['A,a', 'x,y'])), _vary(None))
check('empty header', solve(_vary(['a,', 'x,y'])), _vary(None))
check('reserved header', solve(_vary(['_extra,b', 'x,y'])), _vary(None))
check('short row', solve(_vary(['a,b', '@'])), _vary(None))
check('long row', solve(_vary(['a', 'x,y'])), _vary(None))
check('normal', solve(_vary(['a,b', 'x,y'])), _vary({'a': 'x', 'b': 'y'}))
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| canonical names | {'a': 'cell', 'b': ' y '} | {'a': 'cell', 'b': ' y '} | Passed |
| duplicate canonical | None | None | Passed |
| empty header | None | None | Passed |
| reserved header | None | None | Passed |
| short row | None | None | Passed |
| long row | {'a': 'x'} | None | Failed |
| normal | {'a': 'x', 'b': 'y'} | {'a': 'x', 'b': 'y'} | Passed |
SHA-256 / 29ba6c0c57426b2525e414c8ce3612db59d9f47320063feec671825f1ae78a7c
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
def _vary(value):
if value == '@END': return 3 + 4*N
if isinstance(value, str): return value.replace('@', 'cell' * N)
if isinstance(value, list): return [_vary(x) for x in value]
if isinstance(value, dict): return {_vary(k): _vary(v) for k,v in value.items()}
return value
N = 1
observations = []
def solve(data):
headers = [h.strip().lower() for h in data[0].split(',')]
values = data[1].split(',')
if any(not h for h in headers): return None
if len(set(headers)) != len(headers): return None
if '_extra' in headers: return None
if len(values) > len(headers): return None
return dict(zip(headers, values))
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('canonical names', solve(_vary([' A ,B', '@, y '])), _vary({'a': '@', 'b': ' y '}))
check('duplicate canonical', solve(_vary(['A,a', 'x,y'])), _vary(None))
check('empty header', solve(_vary(['a,', 'x,y'])), _vary(None))
check('reserved header', solve(_vary(['_extra,b', 'x,y'])), _vary(None))
check('short row', solve(_vary(['a,b', '@'])), _vary(None))
check('long row', solve(_vary(['a', 'x,y'])), _vary(None))
check('normal', solve(_vary(['a,b', 'x,y'])), _vary({'a': 'x', 'b': 'y'}))
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| canonical names | {'a': 'cell', 'b': ' y '} | {'a': 'cell', 'b': ' y '} | Passed |
| duplicate canonical | None | None | Passed |
| empty header | None | None | Passed |
| reserved header | None | None | Passed |
| short row | {'a': 'cell'} | None | Failed |
| long row | None | None | Passed |
| normal | {'a': 'x', 'b': 'y'} | {'a': 'x', 'b': 'y'} | Passed |
SHA-256 / 8f9f9f283b70ca92544fb399c4a7dcc3dac4815a4f1da576a9baee1793dd6807
HELD IN THE MEMBER ARCHIVE
The verified repair and its recorded checks are member-only.
This mechanism has 7 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.
Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.
Member access is invitation-based. Sign in with your invited account to inspect the repair.
Sign in to the archive ↗Verification & scope
Deterministic bounded in-memory model. No claim of complete CSV or external format conformance. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:44:40.759715+00:00.
Case digest / 16238eab716e3014a662cad5feeb34c3bd57de7d24f236dcc35196b50e621b94