Model class
3Bconverted WeiboAI/VibeThinker-3BobservedVibeThinker-3B MLX Conversion
This release answers a packaging question, not a training question: can the reasoning model be preserved in a local MLX-compatible form? A small local GSM8K screen scored 40/40, but no conversion-parity suite was recorded and the base has no native tool-calling behavior. The case study therefore reports a useful conversion without inventing a model-quality delta.
Headline Numbers
Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.
Local GSM8K screen
40/40small historical slice; reasoning sanity checkhistoricalTraining delta
Noneformat conversion, not post-trainingderivedNative tool calling
Nonot a drop-in agenthistoricalCompetitive Context
| System | Metric | Score | Size / Class | Comparable? | Readout |
|---|---|---|---|---|---|
| PostTrainLLM VibeThinker-3B MLX | local GSM8K sanity slice | 40/40 | 3B MLX conversion | Direct | Historical local verification of reasoning behavior on a small slice; not a broad benchmark claim. |
| WeiboAI/VibeThinker-3B | upstream reasoning model | source weights | 3B | Not comparable | The conversion derives from this public model, but no controlled pre/post conversion parity table was preserved. |
| Agentic distilled descendant | current tool-calling eval | not recorded | 3B | Not comparable | Public descendant weights exist, but they have no current eval promotion. |
Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.
What this release proves
| Question | Evidence | Decision |
|---|---|---|
| Are unique converted weights public? | Yes | Preserve on Hugging Face |
| Did a local reasoning sanity screen run? | GSM8K 40/40 | Useful, small historical slice |
| Was conversion parity measured? | Not recorded | No parity claim |
| Was the model post-trained by PostTrainLLM? | No | Conversion-only artifact |
| Does it natively call tools? | No | Not an agentic specialist |
Release Blockers
No controlled conversion-parity report
The local sanity result does not prove numerical or benchmark parity with the upstream runtime.
Unblock: Run paired upstream-vs-MLX checks only if a consumer needs this conversion as an active dependency.
No native agentic behavior
The reasoning base does not provide a validated tool-calling interface.
Unblock: Treat it as a reasoning/runtime artifact; evaluate a separately adapted candidate before agent use.