FA-72865 / Probabilistic sketches / Member archive
Bloom filter with double hashing: hash consumes code points instead of UTF-8 bytes · case 05
Non-ASCII keys map to different bits than in the reference writer, producing false negatives across implementations.
Case contract
Input {m, k, add, query}. Each string is hashed with 32-bit FNV-1a over its UTF-8 bytes using seed 0 (h1) and seed 0x5bd1e995 (h2); h2 is forced odd by OR-ing 1. Probe i for i in 0..k-1 is ((h1 + i*h2) mod 2^32) mod m. Return [sorted set bit indices, membership booleans for query] where membership requires every probe bit.
Why this case matters
Bloom filters built with Kirsch-Mitzenmacher double hashing must reproduce the exact same probe sequence across writers and readers; any drift silently changes the bit layout.
One recorded failure
Sample boundary fixtureThis sample comes from the broken implementation of a controlled reproducer.
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| bloom 0 m=66 k=3 | [[0, 6, 8, 14, 15, 29, 32, 33, 37, 45, 58, 61, 64], [true, true, false, false, false, false]] | [[0, 6, 14, 29, 33, 36, 37, 41, 45, 46, 51, 55, 58, 61, 62], [true, true, false, false, false, false]] | Failed |
MEMBER ARCHIVE
The complete case is available to members.
This record includes three runnable implementations, regression fixtures, execution results, and source hashes.
Member access is invitation-based. Sign in with your invited account to inspect the sources.
Sign in to the archive ↗