PostTrainLLM fine-tune report card · schema v1 · compiler 1.0.0
Qwen3-4B ReST Fused
— not recordedThe specialist package format does not record the owner goal that framed the run.
- Target
- Qwen3-4B ReST Fused
- Base model
- Qwen/Qwen3-4B-Instruct-2507 (bf16)
- Candidate
- qwen3-4b-rest-fused
- Method
- teacher-free ReST over checker-passing interleaved trajectories with a file-ops gold depth anchor
- Compiled from
specialists/qwen3-4b-rest-fused(specialist-package)
Decision
Shipped Shipped for a named route only
retain only as a routed file-operations specialist; reject as a general successor and do not use as the Pace default planner
Routing constraint
retain only as a routed file-operations specialist; reject as a general successor and do not use as the Pace default planner. Do not use for: Pace default planner without re-distillation and its product ship gate; unqualified broad general-agent claims.
This candidate is safe only inside that envelope. It is not a general replacement for the base model.
Verification status
Not fully verified. This report does not claim a verified ship. Reasons:
- Primary gate `file_ops_hard_gate` baseline is `historical`, not a current measurement.
- Primary gate `file_ops_hard_gate` candidate is `historical`, not a current measurement.
- Primary gate `file_ops_hard_gate` has no threshold value (state `missing`).
- Primary gate `file_ops_hard_gate` has no passed value (state `missing`).
- Frozen-eval identity is not recorded as a current measurement.
- Train/eval overlap (leakage) was not checked with current evidence.
- Failure reason
- Not applicable
- Failure-reason confidence
- not-applicable
- Lesson
- Not recorded
Next action
Before and after
This candidate does not present an unconditional win: out_of_domain_breadth failed. Target and regression gates are reported independently below.
| Gate | Role | Metric | Baseline | Candidate | Delta | Threshold | Result | n | Frontier ceiling |
|---|---|---|---|---|---|---|---|---|---|
| file_ops_hard_gate | primary | file_ops_hard_gate | 0.7500 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 1.0000 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.2500 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. | — not recordedThe specialist package format records no per-gate threshold. | — not recordedThe specialist package records no ship threshold for the primary gate, so a pass/fail result cannot be derived. | 12 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 1.0000 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). |
| out_of_domain_breadth | breadth | out_of_domain_breadth | 0.6667 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.5556 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | -0.1111 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. | — not recordedThe specialist package format records no per-gate threshold. | no derivedNo threshold was recorded. Derived as failing because the candidate scored below the baseline on a non-primary gate. | 45 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.9778 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). |
Eval identity
| Gate | Suite | Command | Date | Frozen |
|---|---|---|---|---|
| file_ops_hard_gate | file_ops_hard_gate | — not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. | 2026-09-03 | — not recordedThe package does not record whether this suite is frozen. |
| out_of_domain_breadth | out_of_domain_breadth | — not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. | 2026-09-03 | — not recordedThe package does not record whether this suite is frozen. |
Per-slice evidence
No slice evidence was recorded for this candidate.
Cost and performance
| Metric | Value | Source |
|---|---|---|
| Latency | — not recordedTraining duration and normalized per-request latency were not measured. Eval time is candidate depth plus breadth wall time; throughput and peak RSS are the candidate depth values. Per-suite values and raw receipt hashes are preserved in the published result. | specialists/qwen3-4b-rest-fused/eval_report.json#performance.latency_ms |
| RAM / peak RSS | 7755.0781 MB historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | specialists/qwen3-4b-rest-fused/eval_report.json#performance.peak_rss_mb |
| Decode throughput | 13.3970 tok/s historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | specialists/qwen3-4b-rest-fused/eval_report.json#performance.tokens_per_second |
| Training time | — not recordedTraining duration and normalized per-request latency were not measured. Eval time is candidate depth plus breadth wall time; throughput and peak RSS are the candidate depth values. Per-suite values and raw receipt hashes are preserved in the published result. | specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_time_seconds |
| Training cost | 0 USD historicalTeacher-free local ReST iteration; no paid model API was used. Recorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_cost_usd |
| Eval time | 2011.1948 s historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | specialists/qwen3-4b-rest-fused/eval_report.json#performance.eval_time_seconds |
Eval validity and leakage
| Check | Result | Source |
|---|---|---|
| Frontier ceiling | 1.0000 historicalRecorded evidence quality: fresh-paired-frontier-qualified-result-with-hashed-raw-receipts. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | specialists/qwen3-4b-rest-fused/eval_report.json#scores[0].frontier |
| Frozen eval | — not recordedThe package records no frozen held-out split identity. | specialists/qwen3-4b-rest-fused/eval_report.json |
| Train/eval overlap | — not recordedNo train/eval overlap check is recorded for this package. | specialists/qwen3-4b-rest-fused/eval_report.json |
Known eval limitations
- Values come from a committed specialist package, not a canonical factory-run folder: eval commands, dataset hashes, and raw predictions are unavailable for independent replay.
Caveats
- The fresh candidate won file-ops depth 12/12 versus 9/12 and removed all eight stock side effects.
- The same candidate lost frontier-qualified breadth 25/45 versus stock 30/45, so the general-successor decision is reject.
- The tracked result preserves raw receipt hashes; raw traces remain local and gitignored.
- Pace requires a different intent envelope and product-specific ship gate.
- Do not use for: Pace default planner without re-distillation and its product ship gate
- Do not use for: unqualified broad general-agent claims
Source evidence
Every number above traces to one of these artifacts. Content hashes are recorded where the source file is committed.
- model card
specialists/qwen3-4b-rest-fused/model_card.md· committed-package-file · sha256036dce9f2b45629b… - eval report
specialists/qwen3-4b-rest-fused/eval_report.json· committed-package-file · sha256e9924e33ad56792a… - reproducibility lock
specialists/qwen3-4b-rest-fused/tinygpt.lock.json· committed-package-file · sha2569a699a938a99ac6f… - prompt contract
specialists/qwen3-4b-rest-fused/prompt.md· committed-package-file · sha2561571ecc23b132713… - specialist registry entry
specialists/registry.json· committed-registry · sha25634b7b43a62dcece3… - recorded result source
evals/verified-wins/rest-requalification-result-v1.json· historical-record · sha256368104c72ef06c66…· The document the legacy score was recorded in. - published weights
hf://models/posttrainllm/qwen3-4b-rest-fused· external-artifact · Public weight location; not hashed by this compiler.
How to read the evidence states
- measured
- Read directly from a source artifact for this candidate.
- derived
- Computed from other recorded values (for example a delta).
- historical
- Imported from a legacy record without current canonical provenance. Treat as weaker than a measurement.
- skipped
- Deliberately not run for this candidate.
- not recorded
- Evidence should exist but does not. No number is implied.
- not applicable
- The check does not apply to this candidate.