Content-defined chunking for AI corpora
What a gear-hash chunker buys you over fixed blocks and file hashes when shards get re-exported.
Whole-file hashing finds exact duplicates and nothing else. Fixed-size blocks find duplicates only when the shared bytes happen to line up on block boundaries, which they stop doing the moment a single record is inserted near the front of a file. Training corpora are re-exported constantly, with records added, removed and reordered, so both approaches miss most of the sharing that is actually there.
Boundaries from content
A content-defined chunker decides where to cut based on the bytes themselves. Needset uses a gear hash: a rolling hash updated one byte at a time with a lookup table, with a cut declared when the low bits of the hash match a mask. Because the cut depends only on a small window of surrounding bytes, an insertion early in the file shifts the boundaries locally and then the chunker resynchronizes. Everything after the edit produces the same chunks as before.
We clamp chunk sizes with a minimum and maximum so pathological inputs cannot produce a million tiny chunks or one enormous one, and we normalize the mask so the average chunk size is stable across media types.
Determinism is the whole point
The chunker is versioned (gear64-v1) and the table, mask and bounds are fixed for that version. Two nodes chunking the same bytes produce the same chunk hashes, which is what lets the catalog deduplicate across uploads, S3 prefixes and time. It is also what makes plans deterministic: a plan references chunk hashes, so recompiling it against the same catalog reproduces it exactly.
What it finds in practice
In the corpora we have measured, the largest sources of shared bytes were re-exported shards with small edits, evaluation sets duplicated into training buckets, and concatenated files that contained earlier files verbatim. None of these are found by file hashes. All of them are found by content-defined chunking. The savings proof reports them as physical bytes, so you do not have to take our word for it.
See what your corpus is actually worth to your model.
Start free with 50 GB under management. A pilot starts read-only, reports in bytes and dollars, and ends with a number you can audit.