PostTrainLLM fine-tune report card · schema v1 · compiler 1.0.0

Qwen3-4B ReST Fused

not recordedThe specialist package format does not record the owner goal that framed the run.

Target
Qwen3-4B ReST Fused
Base model
Qwen/Qwen3-4B-Instruct-2507 (bf16)
Candidate
qwen3-4b-rest-fused
Method
teacher-free ReST over checker-passing interleaved trajectories with a file-ops gold depth anchor
Compiled from
specialists/qwen3-4b-rest-fused (specialist-package)

Decision

Shipped Shipped for a named route only

retain only as a routed file-operations specialist; reject as a general successor and do not use as the Pace default planner

Routing constraint

retain only as a routed file-operations specialist; reject as a general successor and do not use as the Pace default planner. Do not use for: Pace default planner without re-distillation and its product ship gate; unqualified broad general-agent claims.

This candidate is safe only inside that envelope. It is not a general replacement for the base model.

Verification status

Not fully verified. This report does not claim a verified ship. Reasons:

Failure reason
Not applicable
Failure-reason confidence
not-applicable
Lesson
Not recorded

Next action

not recordedThe specialist package format records no machine-readable next action. The release action lives in the public artifact registry (docs/factory/public-artifacts.md) as prose.

Before and after

This candidate does not present an unconditional win: out_of_domain_breadth failed. Target and regression gates are reported independently below.

Every gate with baseline, candidate, derived delta, threshold, result, sample size, and frontier-ceiling evidence.
Gate Role Metric Baseline Candidate Delta Threshold Result n Frontier ceiling
file_ops_hard_gate primary file_ops_hard_gate 0.7500 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 1.0000 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.2500 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. not recordedThe specialist package format records no per-gate threshold. not recordedThe specialist package records no ship threshold for the primary gate, so a pass/fail result cannot be derived. 12 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 1.0000 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions).
out_of_domain_breadth breadth out_of_domain_breadth 0.6667 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.5556 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). -0.1111 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. not recordedThe specialist package format records no per-gate threshold. no derivedNo threshold was recorded. Derived as failing because the candidate scored below the baseline on a non-primary gate. 45 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.9778 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions).

Eval identity

Which suite produced each gate, how it was invoked, and whether it is frozen.
Gate Suite Command Date Frozen
file_ops_hard_gate file_ops_hard_gate not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. 2026-09-03 not recordedThe package does not record whether this suite is frozen.
out_of_domain_breadth out_of_domain_breadth not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. 2026-09-03 not recordedThe package does not record whether this suite is frozen.

Per-slice evidence

No slice evidence was recorded for this candidate.

Cost and performance

Latency, memory, throughput, and training cost/time. Absent evidence is reported as not measured, never as zero.
MetricValueSource
Latency not recordedTraining duration and normalized per-request latency were not measured. Eval time is candidate depth plus breadth wall time; throughput and peak RSS are the candidate depth values. Per-suite values and raw receipt hashes are preserved in the published result. specialists/qwen3-4b-rest-fused/eval_report.json#performance.latency_ms
RAM / peak RSS 7755.0781 MB historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). specialists/qwen3-4b-rest-fused/eval_report.json#performance.peak_rss_mb
Decode throughput 13.3970 tok/s historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). specialists/qwen3-4b-rest-fused/eval_report.json#performance.tokens_per_second
Training time not recordedTraining duration and normalized per-request latency were not measured. Eval time is candidate depth plus breadth wall time; throughput and peak RSS are the candidate depth values. Per-suite values and raw receipt hashes are preserved in the published result. specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_time_seconds
Training cost 0 USD historicalTeacher-free local ReST iteration; no paid model API was used. Recorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_cost_usd
Eval time 2011.1948 s historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). specialists/qwen3-4b-rest-fused/eval_report.json#performance.eval_time_seconds

Eval validity and leakage

Whether the benchmark is a trustworthy ruler: frontier ceiling, frozen-eval identity, and train/eval overlap.
CheckResultSource
Frontier ceiling 1.0000 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). specialists/qwen3-4b-rest-fused/eval_report.json#scores[0].frontier
Frozen eval not recordedThe package records no frozen held-out split identity. specialists/qwen3-4b-rest-fused/eval_report.json
Train/eval overlap not recordedNo train/eval overlap check is recorded for this package. specialists/qwen3-4b-rest-fused/eval_report.json

Known eval limitations

Caveats

Source evidence

Every number above traces to one of these artifacts. Content hashes are recorded where the source file is committed.

How to read the evidence states

measured
Read directly from a source artifact for this candidate.
derived
Computed from other recorded values (for example a delta).
historical
Imported from a legacy record without current canonical provenance. Treat as weaker than a measurement.
skipped
Deliberately not run for this candidate.
not recorded
Evidence should exist but does not. No number is implied.
not applicable
The check does not apply to this candidate.