Mac runtime benchmarkReport artifact2026-06-06

Huge Preset Decode Throughput

This is a runtime artifact, not a model-quality claim. Its value is operational: local eval loops need throughput, stable serving, and cheap repeated generation.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

Huge decode

696 tok/s96M/Huge preset, ctx 1024historical

Mega pilot

293 tok/s960M pilothistorical

Warm TTFT p99

5.8msreported runtime metrichistorical

Competitive Context

SystemMetricScoreSize / ClassComparable?Readout
posttrainllm Huge presetdecode throughput696 tok/s96MDirectLocal runtime baseline for cheap repeated eval/smoke loops.
posttrainllm Mega pilotdecode throughput293 tok/s960MDirectShows the throughput drop as local model size approaches specialist scale.
External serving stackssame benchmarknot measuredMLX/llama.cpp/Ollama classNot comparableNeeds a shared prompt/config/device table before public competitive serving claims.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Runtime numbers

MetricValueUse
Decode throughput696 tok/sFast local eval/smoke loops
Mega pilot throughput293 tok/sBoundary mapping for larger local models
Warm TTFT p995.8msInteractive serving viability

Release Blockers

Preset-specific

The headline number is not a blanket claim for all HF models or specialists.

Unblock: Attach latency/RAM/tok-s numbers to each future specialist artifact.

Evidence

Next Release Action

Use this as the baseline expectation for future artifact performance tables.