Conversion case studyReport artifact2026-07-03

VibeThinker-3B MLX Conversion

This release answers a packaging question, not a training question: can the reasoning model be preserved in a local MLX-compatible form? A small local GSM8K screen scored 40/40, but no conversion-parity suite was recorded and the base has no native tool-calling behavior. The case study therefore reports a useful conversion without inventing a model-quality delta.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

Model class

3Bconverted WeiboAI/VibeThinker-3Bobserved

Local GSM8K screen

40/40small historical slice; reasoning sanity checkhistorical

Training delta

Noneformat conversion, not post-trainingderived

Native tool calling

Nonot a drop-in agenthistorical

Competitive Context

SystemMetricScoreSize / ClassComparable?Readout
PostTrainLLM VibeThinker-3B MLXlocal GSM8K sanity slice40/403B MLX conversionDirectHistorical local verification of reasoning behavior on a small slice; not a broad benchmark claim.
WeiboAI/VibeThinker-3Bupstream reasoning modelsource weights3BNot comparableThe conversion derives from this public model, but no controlled pre/post conversion parity table was preserved.
Agentic distilled descendantcurrent tool-calling evalnot recorded3BNot comparablePublic descendant weights exist, but they have no current eval promotion.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

What this release proves

QuestionEvidenceDecision
Are unique converted weights public?YesPreserve on Hugging Face
Did a local reasoning sanity screen run?GSM8K 40/40Useful, small historical slice
Was conversion parity measured?Not recordedNo parity claim
Was the model post-trained by PostTrainLLM?NoConversion-only artifact
Does it natively call tools?NoNot an agentic specialist

Release Blockers

No controlled conversion-parity report

The local sanity result does not prove numerical or benchmark parity with the upstream runtime.

Unblock: Run paired upstream-vs-MLX checks only if a consumer needs this conversion as an active dependency.

No native agentic behavior

The reasoning base does not provide a validated tool-calling interface.

Unblock: Treat it as a reasoning/runtime artifact; evaluate a separately adapted candidate before agent use.

Evidence

Next Release Action

Keep this as a conversion and runtime artifact. Add a pinned loader plus parity receipt only when a real Mac-local consumer justifies maintaining it.