FAILURE MAP
← Case archive

FA-72864 / Probabilistic sketches / Member archive

Bloom filter with double hashing: hash consumes code points instead of UTF-8 bytes · case 04

Non-ASCII keys map to different bits than in the reference writer, producing false negatives across implementations.

Member previewVariant 4 · 3 implementations · 7 checks per implementation

Case contract

Input {m, k, add, query}. Each string is hashed with 32-bit FNV-1a over its UTF-8 bytes using seed 0 (h1) and seed 0x5bd1e995 (h2); h2 is forced odd by OR-ing 1. Probe i for i in 0..k-1 is ((h1 + i*h2) mod 2^32) mod m. Return [sorted set bit indices, membership booleans for query] where membership requires every probe bit.

Why this case matters

Bloom filters built with Kirsch-Mitzenmacher double hashing must reproduce the exact same probe sequence across writers and readers; any drift silently changes the bit layout.

One recorded failure

Sample boundary fixture

This sample comes from the broken implementation of a controlled reproducer.

Boundary fixtureActualExpectedOutcome
bloom 0 m=65 k=3[[7, 8, 9, 14, 15, 26, 28, 37, 47, 48, 49, 56, 64], [true, true, false, false, false, false]][[0, 6, 7, 9, 14, 15, 16, 26, 28, 37, 47, 49, 56, 64], [true, true, false, false, false, false]]Failed

MEMBER ARCHIVE

The complete case is available to members.

This record includes three runnable implementations, regression fixtures, execution results, and source hashes.

Member access is invitation-based. Sign in with your invited account to inspect the sources.

Sign in to the archive ↗