49.5M on-device classifierRelease-ready weights2026-07-18

Pace Intent Router v8

This is the most important rejection in the published model set. Pace v8 scored 95.5% on 14,995 source-matched synthetic examples at roughly 3ms, but only 57.1% on a new 63-instance sealed suite that frontier aced. The weights remain useful as a latency floor and as proof that the synthetic generator overfit its template families—not as the production winner.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

Sealed accuracy

57.1%63 held-out instances, two passesmeasured

Frontier retained

57.1%Codex gpt-5.5 scored 100% on the same rulermeasured

Warm latency

3.8msmean local sealed-eval latencymeasured

Source-matched holdout

95.5%14,995 synthetic rows; did not generalizemeasured

Competitive Context

SystemMetricScoreSize / ClassComparable?Readout
Pace Intent Router v8sealed exact intent accuracy57.1%49.5MDirectFastest local entry, but rejected because fresh phrasing exposed distribution overfit.
Qwen3-4B-Instruct-2507 4-bitsame sealed exact intent accuracy93.7%4BDirectCapability winner among local entries at 211ms mean warm latency.
Apple FoundationModelssame sealed exact intent accuracy92.1%on-device system modelDirect522ms mean warm latency; much slower than Pace v8 but materially more accurate on fresh phrasing.
Codex gpt-5.5same sealed exact intent accuracy100%frontier anchorDirectPassed the frozen 99% frontier-ceiling gate. Its amortized batch latency is not comparable with per-instance local timing.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Attempts and decisions

Evidence stageAccuracyLatencyDecisionLesson
Source-matched synthetic holdout95.5%3.1ms p50promisingThe pipeline learned its generator distribution
Sealed V157.1%3.8ms meanreject-production-winnerFresh user-like phrasing broke the apparent win

Sealed V1 result

EntryAccuracyUnknown recallWarm latencyReadout
Codex gpt-5.5100.0%100.0%519ms amortized batchFrontier ruler passed
Qwen3 4B 4-bit93.7%77.8%211msLocal capability winner
Apple FoundationModels92.1%77.8%522msAccurate, slower
Pace v857.1%55.6%3.8msLatency floor; rejected

Release Blockers

Fresh-distribution generalization failed

The earlier 95.5% result was source-matched to the synthetic generator. Accuracy fell to 57.1% on leakage-checked sealed V1.

Unblock: Generate public-development examples from failure themes—not sealed prompts—and require a newly generated sealed V2 for any successor claim.

Rejected as the production winner

Qwen3 4B and Apple FoundationModels both exceeded 92% on the same sealed suite.

Unblock: Keep v8 only as a latency floor and training-pipeline artifact until a successor clears the sealed gate.

Evidence

Next Release Action

Do not tune on sealed V1. Build a broader public-development corpus from the failure themes, train a successor, and judge it once on a newly generated sealed V2.