Browser performance artifactReport artifact2026-09-02

Browser WebGPU Training Speedup

The browser playground now has a retained, adapter-qualified performance win for one configuration. Two WASM and two WebGPU runs used the same Large preset, corpus, seed, and 20-step budget. This qualifies Large on the measured M5 Pro; the older Small/Medium/XL curve remains historical until separately reproduced.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

WebGPU speedup

10.67xmeasured · Large · Apple M5 Promeasured

Loss drift

4.72%maximum across two paired comparisonsmeasured

Hardware gate

passApple · metal-3 · no fallbackmeasured

Competitive Context

SystemMetricScoreSize / ClassComparable?Readout
posttrainllm WebGPUtraining step speedup10.67xLarge · d_model=192 · 20 stepsDirectMedian 137.1 ms/step across two WebGPU runs on the same Apple hardware and frozen inputs.
posttrainllm WASM SIMDtraining step speedup1.0xsame browser model/configDirectMedian 1462.5 ms/step across the two matching WASM runs.
Native Mac runtimesbrowser training benchmarknot measuredMLX/Metal classNot comparableNative runtimes are the right competition for production throughput, but not for the browser-learning artifact.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Performance readout

VariantResultInterpretation
WASM · Large1462.5 ms/step medianTwo matching runs
WebGPU · Large137.1 ms/step median10.67x measured speedup
Final-loss drift4.72% maximumPasses the frozen <5% gate
Other presets2.6x–12.1x historicalStill unqualified

Release Blockers

Evidence

Next Release Action

Keep Large as the measured M5 Pro win; reproduce Small, Medium, and XL independently before promoting the historical curve.