FA-48446 / Delimited text / Open access
A text checksum is calculated from UTF-8 bytes instead of code points · case 01
A structured table violates the declared record or column contract.
ROOT CAUSE
A text checksum is calculated from UTF-8 bytes instead of code points.
THE FAILURE
A text checksum is calculated from UTF-8 bytes instead of code points.
Unsuccessful approach: The alternate implementation still violates the same declared invariant: a text checksum is calculated from utf-8 bytes instead of code points.
Case contract
Decode rows encoded as comma cells followed by |decimal check. The check is the sum of Unicode code points in the exact comma payload modulo 97; this is only a corruption-detection toy, not cryptographic integrity. Use the last pipe as trailer separator, permitting pipes in cells. Reject malformed or mismatched checks. Return fields.
Why this case matters
Delimited interchange needs explicit framing, schema and field semantics at ingestion and emission boundaries.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
def _vary(value):
if value == '@END': return 3 + 4*N
if isinstance(value, str): return value.replace('@', 'cell' * N)
if isinstance(value, list): return [_vary(x) for x in value]
if isinstance(value, dict): return {_vary(k): _vary(v) for k,v in value.items()}
return value
N = 1
observations = []
def solve(data):
parts=data.rsplit('|',1)
if len(parts)!=2: return None
payload,token=parts
if not token.isascii() or not token.isdecimal(): return None
check=int(token)
if check>=97: return None
expected=sum(payload.encode('utf-8'))%97
if check!=expected: return None
return payload.split(',')
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('comma', solve(_vary('a,b|45')), _vary(['a', 'b']))
check('empty', solve(_vary('|0')), _vary(['']))
check('pipe data', solve(_vary('a|b|28')), _vary(['a|b']))
check('unicode', solve(_vary('é|39')), _vary(['é']))
check('astral checksum', solve(_vary('😀|84')), _vary(['😀']))
check('bad checksum', solve(_vary('a|1')), _vary(None))
check('bad trailer', solve(_vary('a|-1')), _vary(None))
check('unreduced', solve(_vary('a|97')), _vary(None))
check('normal', solve(_vary('z|25')), _vary(['z']))
check('variant checksum width', solve(('x'*N)+'|'+str((120*N)%97)), ['x'*N])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| comma | ['a', 'b'] | ['a', 'b'] | Passed |
| empty | [''] | [''] | Passed |
| pipe data | ['a|b'] | ['a|b'] | Passed |
| unicode | None | ['é'] | Failed |
| astral checksum | None | ['😀'] | Failed |
| bad checksum | None | None | Passed |
| bad trailer | None | None | Passed |
| unreduced | None | None | Passed |
| normal | ['z'] | ['z'] | Passed |
| variant checksum width | ['x'] | ['x'] | Passed |
SHA-256 / 1ef1bf4827e641f357f71b91a0d80d4092fc028ede43511ae6465320c2aa9702
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
def _vary(value):
if value == '@END': return 3 + 4*N
if isinstance(value, str): return value.replace('@', 'cell' * N)
if isinstance(value, list): return [_vary(x) for x in value]
if isinstance(value, dict): return {_vary(k): _vary(v) for k,v in value.items()}
return value
N = 1
observations = []
def solve(data):
parts=data.rsplit('|',1)
if len(parts)!=2: return None
payload,token=parts
if not token.isascii() or not token.isdecimal(): return None
check=int(token)
if check>=97: return None
expected=sum(payload.encode('utf-16-le'))%97
if check!=expected: return None
return payload.split(',')
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
check('comma', solve(_vary('a,b|45')), _vary(['a', 'b']))
check('empty', solve(_vary('|0')), _vary(['']))
check('pipe data', solve(_vary('a|b|28')), _vary(['a|b']))
check('unicode', solve(_vary('é|39')), _vary(['é']))
check('astral checksum', solve(_vary('😀|84')), _vary(['😀']))
check('bad checksum', solve(_vary('a|1')), _vary(None))
check('bad trailer', solve(_vary('a|-1')), _vary(None))
check('unreduced', solve(_vary('a|97')), _vary(None))
check('normal', solve(_vary('z|25')), _vary(['z']))
check('variant checksum width', solve(('x'*N)+'|'+str((120*N)%97)), ['x'*N])
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| comma | ['a', 'b'] | ['a', 'b'] | Passed |
| empty | [''] | [''] | Passed |
| pipe data | ['a|b'] | ['a|b'] | Passed |
| unicode | ['é'] | ['é'] | Passed |
| astral checksum | None | ['😀'] | Failed |
| bad checksum | None | None | Passed |
| bad trailer | None | None | Passed |
| unreduced | None | None | Passed |
| normal | ['z'] | ['z'] | Passed |
| variant checksum width | ['x'] | ['x'] | Passed |
SHA-256 / cc4d10a778b399e048c1e5d12030fac1d9687d826d9fa47b58b0a9827fef73c3
HELD IN THE MEMBER ARCHIVE
The verified repair and its recorded checks are member-only.
This mechanism has 10 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.
Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.
Member access is invitation-based. Sign in with your invited account to inspect the repair.
Sign in to the archive ↗Verification & scope
Deterministic bounded in-memory model. No claim of complete CSV or external format conformance. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:44:50.879028+00:00.
Case digest / bcdc59d9ed19639b419b7dd60582a858c582b70dc9fd7b2fb9b5f11016b57bc0