A failed model that won us an answerReport artifact2026-09-03

Needle 2: The 45M Capacity Boundary

The lab first gave Needle the oracle task family, then trained plain and distractor-negative data under standard and safety-weighted losses. Every arm passed tiny overfit, but the best successor scored 26/94 versus stock's 32/94 and every arm produced ten destructive-action bypasses. The safety stop prevented a wasteful sweep and kept sealed V2 unopened. That is a useful capacity-boundary result, not a product candidate.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

Tiny-overfit gates

4/4all arms reached exact 1.0measured

Best successor

26/9427.7% vs stock 34.0%measured

Destructive bypasses

10 vs 2best arm vs stock; unsafe stopmeasured

Competitive Context

SystemMetricScoreSize / ClassComparable?Readout
Needle 2 full catalogexact tool selection34.0%94 public fixturesDirectIncluded eight out-of-scope false calls and two destructive-action calls.
Needle 2 oracle task catalogexact tool selection38.3%same 94 fixturesDirectThe upstream task family was supplied; this does not prove Needle can discover it.
Best trained 45M successorexact tool selection27.7%94 public-development fixturesDirectDistractor plus safety weighting was the best arm, but remained below stock and failed the destructive-action stop.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Successor factorial

Training dataLossExactDestructive bypasses
PlainStandard24/9410
PlainSafety24/9410
DistractorStandard25/9410
DistractorSafety26/9410

Release Blockers

45M capacity boundary

All four trained arms passed memorization but regressed public exactness and multiplied destructive bypasses.

Unblock: Open a fresh scoped experiment for a 1.7B successor under the same frozen public and sealed gates.

Evidence

Next Release Action

Close the 45M recipe-search lane. Keep the complete factorial as a hands-on learning artifact; any 1.7B successor is a fresh experiment.