FA-79651 / Barcode symbology encoding / Open access
Byte mode counts characters instead of UTF-8 bytes · case 01
Payloads with accented letters or the euro sign overflow the chosen version and are truncated by the encoder.
ROOT CAUSE
The byte segment length is the number of code points.
VERIFIED REPAIR
Count the UTF-8 encoded bytes.
Unsuccessful approach: Encoding as Latin-1 with replacement still yields one byte per character.
Case contract
Compute the length in bits of one QR segment: 4-bit mode indicator, a character count indicator whose width depends on mode and version band (1-9, 10-26, 27-40): numeric 10/12/14, alphanumeric 9/11/13, byte 8/16/16, and the data bits: numeric 10 bits per 3 digits plus 4 or 7 for a remainder of 1 or 2, alphanumeric 11 bits per pair plus 6 for a single, byte 8 bits per UTF-8 byte. Alphanumeric allows 0-9 A-Z space $ % * + - . / :. A count that does not fit the indicator is too-long. Errors: version, charset, mode, too-long.
Why this case matters
Retail, logistics, pharmacy and document workflows depend on encoders that produce exactly the module pattern, code-set switches, separators and quiet zones scanners expect; one misplaced module or separator makes a label unreadable or, worse, scan as different data.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(mode, data, version):
if not 1 <= version <= 40:
return {'error': 'version'}
band = 0 if version <= 9 else (1 if version <= 26 else 2)
ALNUM = '0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ $%*+-./:'
if mode == 'numeric':
if not all(c in '0123456789' for c in data):
return {'error': 'charset'}
count = len(data)
bits = 10 * (count // 3) + [0, 4, 7][count % 3]
cci = [10, 12, 14][band]
elif mode == 'alnum':
if not all(c in ALNUM for c in data):
return {'error': 'charset'}
count = len(data)
bits = 11 * (count // 2) + 6 * (count % 2)
cci = [9, 11, 13][band]
elif mode == 'byte':
count = len(data)
bits = 8 * count
cci = [8, 16, 16][band]
else:
return {'error': 'mode'}
if count >= 1 << cci:
return {'error': 'too-long'}
return 4 + cci + bits
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[[('byte', '1€11€é', 40), 108], [('byte', '€éa1a1 ', 9), 92], [('alnum', '-/$é+/%é+', 0), {'error': 'version'}], [('alnum', '-', 1), 19], [('alnum', '%', 1), 19], [('alnum', 'ééA- q', 9), {'error': 'charset'}], [('byte', '', 1), 12], [('byte', '€1é ', 27), 76]], [[('byte', 'ébba', 27), 60], [('byte', 'é1 ', 27), 52], [('numeric', '053142', 0), {'error': 'version'}], [('numeric', '089405', 10), 36], [('kanji', 'b1ab11', 5), {'error': 'mode'}], [('alnum', 'q+$ ', 40), {'error': 'charset'}], [('alnum', '++$:é1*.+', 10), {'error': 'charset'}], [('byte', 'ébé', 1), 52]], [[('byte', '1abé', 1), 52], [('byte', 'é ', 5), 44], [('numeric', '00575', 9), 31], [('numeric', '37219', 5), 31], [('numeric', '', 40), 18], [('numeric', '716925614701', 5), 54], [('kanji', 'é b€b', 10), {'error': 'mode'}], [('byte', '€a€', 10), 76]], [[('byte', 'é€ééa ba', 15), 124], [('byte', 'b€a', 1), 52], [('byte', '€', 41), {'error': 'version'}], [('numeric', '', 0), {'error': 'version'}], [('numeric', '2', 1), 18], [('alnum', '%1.qZ-', 5), {'error': 'charset'}], [('alnum', 'A:*éZA', 40), {'error': 'charset'}], [('byte', 'éb aa€a', 10), 100]], [[('byte', ' a bbé', 10), 76], [('byte', '111bé1', 5), 68], [('kanji', 'éb', 1), {'error': 'mode'}], [('alnum', '.', 5), 19], [('numeric', '4813206244', 15), 50], [('numeric', '19928871', 5), 41], [('numeric', '', 1), 14], [('byte', 'a€', 26), 52]]]
labels = ["regression: byte mode length unit", "repair trap", "combined fault", "control", "control", "boundary", "boundary", "control"]
for i, (args, expected) in enumerate(fixtures[N-1]):
check("%s %d" % (labels[i % len(labels)], i), solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| regression: byte mode length unit 0 | 68 | 108 | Failed |
| repair trap 1 | 68 | 92 | Failed |
| combined fault 2 | {'error': 'version'} | {'error': 'version'} | Passed |
| control 3 | 19 | 19 | Passed |
| control 4 | 19 | 19 | Passed |
| boundary 5 | {'error': 'charset'} | {'error': 'charset'} | Passed |
| boundary 6 | 12 | 12 | Passed |
| control 7 | 52 | 76 | Failed |
SHA-256 / 196e6602fcf7ec1a21c4080008c67c1b948764b3b6aac67ae842c8a74f6c4ac1
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(mode, data, version):
if not 1 <= version <= 40:
return {'error': 'version'}
band = 0 if version <= 9 else (1 if version <= 26 else 2)
ALNUM = '0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ $%*+-./:'
if mode == 'numeric':
if not all(c in '0123456789' for c in data):
return {'error': 'charset'}
count = len(data)
bits = 10 * (count // 3) + [0, 4, 7][count % 3]
cci = [10, 12, 14][band]
elif mode == 'alnum':
if not all(c in ALNUM for c in data):
return {'error': 'charset'}
count = len(data)
bits = 11 * (count // 2) + 6 * (count % 2)
cci = [9, 11, 13][band]
elif mode == 'byte':
count = len(data.encode('latin-1', 'replace'))
bits = 8 * count
cci = [8, 16, 16][band]
else:
return {'error': 'mode'}
if count >= 1 << cci:
return {'error': 'too-long'}
return 4 + cci + bits
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[[('byte', '1€11€é', 40), 108], [('byte', '€éa1a1 ', 9), 92], [('alnum', '-/$é+/%é+', 0), {'error': 'version'}], [('alnum', '-', 1), 19], [('alnum', '%', 1), 19], [('alnum', 'ééA- q', 9), {'error': 'charset'}], [('byte', '', 1), 12], [('byte', '€1é ', 27), 76]], [[('byte', 'ébba', 27), 60], [('byte', 'é1 ', 27), 52], [('numeric', '053142', 0), {'error': 'version'}], [('numeric', '089405', 10), 36], [('kanji', 'b1ab11', 5), {'error': 'mode'}], [('alnum', 'q+$ ', 40), {'error': 'charset'}], [('alnum', '++$:é1*.+', 10), {'error': 'charset'}], [('byte', 'ébé', 1), 52]], [[('byte', '1abé', 1), 52], [('byte', 'é ', 5), 44], [('numeric', '00575', 9), 31], [('numeric', '37219', 5), 31], [('numeric', '', 40), 18], [('numeric', '716925614701', 5), 54], [('kanji', 'é b€b', 10), {'error': 'mode'}], [('byte', '€a€', 10), 76]], [[('byte', 'é€ééa ba', 15), 124], [('byte', 'b€a', 1), 52], [('byte', '€', 41), {'error': 'version'}], [('numeric', '', 0), {'error': 'version'}], [('numeric', '2', 1), 18], [('alnum', '%1.qZ-', 5), {'error': 'charset'}], [('alnum', 'A:*éZA', 40), {'error': 'charset'}], [('byte', 'éb aa€a', 10), 100]], [[('byte', ' a bbé', 10), 76], [('byte', '111bé1', 5), 68], [('kanji', 'éb', 1), {'error': 'mode'}], [('alnum', '.', 5), 19], [('numeric', '4813206244', 15), 50], [('numeric', '19928871', 5), 41], [('numeric', '', 1), 14], [('byte', 'a€', 26), 52]]]
labels = ["regression: byte mode length unit", "repair trap", "combined fault", "control", "control", "boundary", "boundary", "control"]
for i, (args, expected) in enumerate(fixtures[N-1]):
check("%s %d" % (labels[i % len(labels)], i), solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| regression: byte mode length unit 0 | 68 | 108 | Failed |
| repair trap 1 | 68 | 92 | Failed |
| combined fault 2 | {'error': 'version'} | {'error': 'version'} | Passed |
| control 3 | 19 | 19 | Passed |
| control 4 | 19 | 19 | Passed |
| boundary 5 | {'error': 'charset'} | {'error': 'charset'} | Passed |
| boundary 6 | 12 | 12 | Passed |
| control 7 | 52 | 76 | Failed |
SHA-256 / c3a36e978961193cf57c2460520f605176ef39d4f4c33f4ec7e2aa88f832de54
3 / The verified repair
Exit 0"""Failure Map reference implementation. Python standard library only."""
import json
N = 1
observations = []
def solve(mode, data, version):
if not 1 <= version <= 40:
return {'error': 'version'}
band = 0 if version <= 9 else (1 if version <= 26 else 2)
ALNUM = '0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ $%*+-./:'
if mode == 'numeric':
if not all(c in '0123456789' for c in data):
return {'error': 'charset'}
count = len(data)
bits = 10 * (count // 3) + [0, 4, 7][count % 3]
cci = [10, 12, 14][band]
elif mode == 'alnum':
if not all(c in ALNUM for c in data):
return {'error': 'charset'}
count = len(data)
bits = 11 * (count // 2) + 6 * (count % 2)
cci = [9, 11, 13][band]
elif mode == 'byte':
count = len(data.encode('utf-8'))
bits = 8 * count
cci = [8, 16, 16][band]
else:
return {'error': 'mode'}
if count >= 1 << cci:
return {'error': 'too-long'}
return 4 + cci + bits
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[[('byte', '1€11€é', 40), 108], [('byte', '€éa1a1 ', 9), 92], [('alnum', '-/$é+/%é+', 0), {'error': 'version'}], [('alnum', '-', 1), 19], [('alnum', '%', 1), 19], [('alnum', 'ééA- q', 9), {'error': 'charset'}], [('byte', '', 1), 12], [('byte', '€1é ', 27), 76]], [[('byte', 'ébba', 27), 60], [('byte', 'é1 ', 27), 52], [('numeric', '053142', 0), {'error': 'version'}], [('numeric', '089405', 10), 36], [('kanji', 'b1ab11', 5), {'error': 'mode'}], [('alnum', 'q+$ ', 40), {'error': 'charset'}], [('alnum', '++$:é1*.+', 10), {'error': 'charset'}], [('byte', 'ébé', 1), 52]], [[('byte', '1abé', 1), 52], [('byte', 'é ', 5), 44], [('numeric', '00575', 9), 31], [('numeric', '37219', 5), 31], [('numeric', '', 40), 18], [('numeric', '716925614701', 5), 54], [('kanji', 'é b€b', 10), {'error': 'mode'}], [('byte', '€a€', 10), 76]], [[('byte', 'é€ééa ba', 15), 124], [('byte', 'b€a', 1), 52], [('byte', '€', 41), {'error': 'version'}], [('numeric', '', 0), {'error': 'version'}], [('numeric', '2', 1), 18], [('alnum', '%1.qZ-', 5), {'error': 'charset'}], [('alnum', 'A:*éZA', 40), {'error': 'charset'}], [('byte', 'éb aa€a', 10), 100]], [[('byte', ' a bbé', 10), 76], [('byte', '111bé1', 5), 68], [('kanji', 'éb', 1), {'error': 'mode'}], [('alnum', '.', 5), 19], [('numeric', '4813206244', 15), 50], [('numeric', '19928871', 5), 41], [('numeric', '', 1), 14], [('byte', 'a€', 26), 52]]]
labels = ["regression: byte mode length unit", "repair trap", "combined fault", "control", "control", "boundary", "boundary", "control"]
for i, (args, expected) in enumerate(fixtures[N-1]):
check("%s %d" % (labels[i % len(labels)], i), solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| regression: byte mode length unit 0 | 108 | 108 | Passed |
| repair trap 1 | 92 | 92 | Passed |
| combined fault 2 | {'error': 'version'} | {'error': 'version'} | Passed |
| control 3 | 19 | 19 | Passed |
| control 4 | 19 | 19 | Passed |
| boundary 5 | {'error': 'charset'} | {'error': 'charset'} | Passed |
| boundary 6 | 12 | 12 | Passed |
| control 7 | 76 | 76 | Passed |
SHA-256 / 0d2e89a3760a0ee1d3d1d6427b8465eb04f07e85bafe5f2371b07db55403b3c4
Verification & scope
A deterministic bounded teaching model with a stipulated contract; it makes no claim of conformance to any published specification. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:49:46.387038+00:00.
Case digest / 9ee2f92dab2c512d6b82555b8677a270f47cf9ba317ec87507fbf758191b3245