The control plane between your data and your models.
Needset is a single service with a catalog, a compiler, a proof engine and an operator portal. It runs in your VPC or ours, talks to any S3-compatible store, and never needs to own your originals.
Catalog
One deduplicated index of everything your models could learn from.
Chunk-level dedup
Files are split with a content-defined chunker so that shared bytes are found even when files are re-exported, reordered or concatenated. Shared blocks are stored once and referenced by hash.
S3 prefix registration
Register a bucket prefix and Needset scans it in place. Originals stay where they are. Materializations stream from your store through a spooled view when a job needs them.
Streaming ingest
Multipart upload streams straight to the catalog with a single-pass tar bundle option for large trees. Ingest is memory-mapped, so a 40 GB shard does not need 40 GB of RAM.
Codec selection
Each asset gets a deterministic codec choice based on measured compressibility, so decode cost is known before a plan is compiled.
Lineage
Every asset carries source, version, media type and the block overlap index with every other asset. Evaluation sets can be fenced out of training selections by lineage, not by naming convention.
Formats
Works on bytes, so it is format-agnostic: parquet, jsonl, arrow, webdataset tars, images, audio, checkpoints. Semantic candidates are an optional layer on top.
Compiler
Turn demand into the smallest data set that satisfies it, inside a budget.
Demand from evidence
Import lm-eval-harness results or record batch failures directly. Demand is derived from where the model is actually weak, with a configurable half-life so stale failures decay.
Budgets and floors
Set a byte budget and a decode-time budget. Coverage floors per need are hard constraints, not soft weights, so the plan is either feasible or it tells you exactly which floor cannot be met.
Exact lazy greedy
Selection is a single deterministic pass over an overlap-aware state with a lazy-evaluation heap. Shared bytes between candidates are counted once, which is what makes physical savings real.
Explain
Every plan has an explanation: why each asset was selected or rejected, which need it covers, what it cost in physical bytes, and which constraint bound the result.
Verify
Verification is recompilation. If the catalog and inputs are the same, the plan hash is the same. Plans can also be HMAC-signed so a downstream job can trust one without a database round-trip.
Quality and trust gates
Minimum quality and trust scores, per-source caps and model-version filters keep a plan inside policy without hand-curating lists.
Pilots, proofs and economics
Turn a plan into a number finance will sign.
Governed pilots
A pilot binds a plan to an eligibility check, an approver and a window. Approval requires an authenticated principal; in production that means SSO.
Savings proof
Reports physical bytes avoided (storage, transfer, decode) and dollars at the unit costs you enter. Signed, reproducible, and exportable as CSV for reconciliation.
Metering
Prometheus endpoint and CSV export of metered incremental deduped bytes, so the invoice and your own dashboards agree.
Operate it
Everything is available three ways: portal, REST API and Python SDK.
Operator portal
Overview, data assets, demand, plans with explain view, pilots, economics, audit, automation and service accounts. Strict CSP with per-load nonce, SSO sign-in.
Automation
Schedules recompile plans on a cadence. HMAC-signed webhooks fire on plan, pilot, proof and anchor events so your orchestrator can react.
SDK and API
Standard-library-only Python client, so it installs anywhere. Every portal action is an API call you can script.
See what your corpus is actually worth to your model.
Start free with 50 GB under management. A pilot starts read-only, reports in bytes and dollars, and ends with a number you can audit.