Scan. Register. Compile. Prove.
A Needset engagement is designed so the first decision costs nothing and every later one is backed by a number you can reproduce.
Scan locally
Run Needset Scout on a representative directory, or use the in-browser scan on the home page. Both work entirely on your machine: exact duplicate ratio, compressed size and a preliminary storage estimate. No account required.
Register assets
Create a free workspace, connect your bucket under Storage (read-only credentials, encrypted at rest) or upload a sample. Needset chunks each asset with a deterministic content-defined chunker, deduplicates at block level and records lineage, media type and a codec choice.
Describe demand
Import evaluation results (lm-eval-harness is supported directly) or record batch failures with the SDK. Each becomes a weighted need. You can override need weights explicitly when you already know what the model is missing.
Compile a plan
Set a byte budget, a decode-time budget, quality and trust minimums, optional coverage floors and per-source caps. The compiler returns a deterministic plan with a hash, a feasibility verdict and a full explanation of selections and rejections.
Verify
Anyone with a viewer key can verify a plan: Needset recompiles it and compares hashes. Signed plans can also be verified offline with the tenant's plan-signing key.
Pilot
Bind the plan to a pilot with an approver and a window. Materialize only the selected block set for the pilot job. Run it side by side with the existing path.
Prove and meter
Enter your unit costs. The savings proof reports physical bytes avoided and dollars, signed. Metering runs continuously on incremental deduped bytes and is exposed via Prometheus and CSV.
The same flow, as code
from needset.client import Needset
ns = Needset("https://app.needset.ai", api_key=OPERATOR_KEY)
# register an S3 prefix in place (read-only)
ns.data.scan(prefix="shards/2026-09/")
# demand from evidence
ns.failures.import_lm_eval("results/lm-eval.json", model_version="v12")
# compile under budgets, with coverage floors as hard constraints
plan = ns.plans.compile(
byte_budget=1_000_000_000,
decode_ms_budget=5_000,
min_quality=0.6,
min_coverage={"code": 0.3, "math": 0.2},
)
print(plan["plan_id"], plan["feasible"], plan["selected_bytes"])
why = ns.plans.explain(plan["plan_id"])
pilot = ns.pilots.create("pilot-2026-09", plan["plan_id"])
proof = ns.economics.prove(plan_id=plan["plan_id"])
Field names abbreviated for readability. See Docs.
Ready to see the flow on your own data?
Assessments are fixed price and read-only. Most teams have a first plan within a week of granting prefix access.