Missing-evidence case studyBlocked2026-07-03

VibeThinker-3B Agentic Distilled

The experiment asked whether a reasoning-strong 3B base could become a better tool-calling student than Qwen3-4B. The fused weights were published, but the repository does not preserve a current before/after eval, regression review, or ship decision. This page is the honest case study of an unfinished proof: the artifact exists; the conclusion does not.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

Public weights

YesHugging Face model repositoryobserved

Current agentic eval

Missingno promoted before/after resultnot-measured

Reasoning retention

Unknownno post-distillation regression gatenot-measured

Decision

Inconclusiveweights are not a ship decisionderived

Competitive Context

SystemMetricScoreSize / ClassComparable?Readout
VibeThinker-3B agentic distilledcurrent tool-calling gatenot recorded3BDirectThe intended candidate exists, but the required result is absent.
VibeThinker-3B MLX basehistorical GSM8K sanity slice40/403BDirectionalShows the source base's reasoning signal, not retained reasoning after agentic distillation.
Qwen3-4B ReSTrecorded file-ops depth / breadth100% / 65%4BNot comparableThe measured incumbent research candidate; no matched VibeThinker result exists.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Evidence audit

Required case-study elementStateWhat can be said
Base identityRecordedVibeThinker-3B MLX reasoning base
Distilled fused weightsPublicArtifact is preserved
Frozen baselineMissingNo before number
Current candidate evalMissingNo after number
Breadth/reasoning regressionMissingRetention unknown
Ship decisionAbsentInconclusive archive only

Release Blockers

No current before/after eval

Public weights alone cannot establish whether tool-calling improved or whether reasoning regressed.

Unblock: Freeze a matched baseline, agentic primary gate, and reasoning/breadth regression gate; then run both once under the same harness.

No specialist package or routing decision

The repository has no canonical package metadata that defines safe use or a ship/reject outcome.

Unblock: Only package after the frozen eval produces a defensible decision.

Evidence

Next Release Action

Leave the weights public but unpromoted. If this lineage becomes active again, start with an evidence-only baseline/candidate evaluation—not more training.