WebGPU speedup
10.67xmeasured · Large · Apple M5 PromeasuredBrowser WebGPU Training Speedup
The browser playground now has a retained, adapter-qualified performance win for one configuration. Two WASM and two WebGPU runs used the same Large preset, corpus, seed, and 20-step budget. This qualifies Large on the measured M5 Pro; the older Small/Medium/XL curve remains historical until separately reproduced.
Headline Numbers
Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.
Loss drift
4.72%maximum across two paired comparisonsmeasuredHardware gate
passApple · metal-3 · no fallbackmeasuredCompetitive Context
| System | Metric | Score | Size / Class | Comparable? | Readout |
|---|---|---|---|---|---|
| posttrainllm WebGPU | training step speedup | 10.67x | Large · d_model=192 · 20 steps | Direct | Median 137.1 ms/step across two WebGPU runs on the same Apple hardware and frozen inputs. |
| posttrainllm WASM SIMD | training step speedup | 1.0x | same browser model/config | Direct | Median 1462.5 ms/step across the two matching WASM runs. |
| Native Mac runtimes | browser training benchmark | not measured | MLX/Metal class | Not comparable | Native runtimes are the right competition for production throughput, but not for the browser-learning artifact. |
Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.
Performance readout
| Variant | Result | Interpretation |
|---|---|---|
| WASM · Large | 1462.5 ms/step median | Two matching runs |
| WebGPU · Large | 137.1 ms/step median | 10.67x measured speedup |
| Final-loss drift | 4.72% maximum | Passes the frozen <5% gate |
| Other presets | 2.6x–12.1x historical | Still unqualified |