From model failures to a data working set
How the compiler turns evaluation results into a budgeted, explainable selection.
Training data selection is usually done by hand: someone decides that the model is weak at code, finds some code, and adds it. The Needset compiler makes that decision from evidence, under a budget, and writes down why.
Demand
Demand is a vector of weighted needs. It is built from recorded failures (a batch of unit tests the model failed, tagged code) and imported evaluation results (an lm-eval-harness run where gsm8k dropped). Each item has a weight and a timestamp, and demand decays with a configurable half-life so a failure from six months ago does not dominate today's plan. You can override the vector explicitly when you already know what you want.
Supply
Every cataloged asset has a need coverage profile, a physical byte cost (its unique chunks, not its logical size), a decode-time estimate from its codec, and quality and trust scores. Assets overlap: two shards that share 60% of their chunks do not cost twice as much to select together. The compiler tracks that with a block-overlap index.
Selection
The compiler runs a single deterministic pass over a selection state. Coverage floors are hard constraints: if the plan must cover 30% of the code need, the plan is infeasible until it does, and the explanation says which floor could not be met within budget. Beyond the floors, it is an exact lazy-greedy: candidates sit in a heap keyed by marginal demand covered per marginal physical byte, and a candidate is only re-scored when it reaches the top, because marginal gains can only decrease as others are selected. Shared bytes are charged once.
Explanation
The output is a plan with a hash and a full explanation: for each asset, selected or rejected, the need it covers, its marginal physical cost, and the constraint that decided it. Rejected assets say why: over budget, below quality, capped per source, or dominated by an asset that covered the same need more cheaply.
Verification
Because everything above is deterministic, verification is recompilation. Same catalog, same demand, same budgets: same hash. That is a stronger guarantee than a signature, and we offer both.
See what your corpus is actually worth to your model.
Start free with 50 GB under management. A pilot starts read-only, reports in bytes and dollars, and ends with a number you can audit.