Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-31 14:30:10 -07:00
parent 4ed4763557
commit 3d35f4953f
5 changed files with 248 additions and 15 deletions
+54 -1
View File
@@ -331,7 +331,8 @@ To continue a completed run without discarding its trained task heads, pass its
Continuation restores the backbone and all four heads strictly, then starts a fresh
optimizer and learning-rate schedule; `--model` remains reserved for an untrained local
upstream checkpoint. The first base run was still improving when its three-epoch schedule
ended, so continue its selected checkpoint conservatively before changing architecture:
ended, so its selected checkpoint was continued conservatively before changing
architecture:
```bash
ml/purpose-classifier/venv/bin/python -u \
@@ -350,6 +351,58 @@ ml/purpose-classifier/venv/bin/python -u \
--overwrite-output
```
That continuation reached 80.31% primary and 78.28% hard-slice validation accuracy;
calibrated mixed F1 reached 67.12%. It remained far below purpose-lite, while primary
training loss and validation accuracy were still improving. Do not chain another plain
continuation. The bounded next experiment distills the mature purpose-lite boundary
teacher into the continued deep checkpoint while retaining direct primary labels and all
three auxiliary losses.
Create a teacher cache bound to the history-augmented split:
```bash
ml/purpose-classifier/venv/bin/python -u \
ml/purpose-classifier/cache_teacher.py \
--dataset-dir \
ml/purpose-classifier/.artifacts/dataset-v1-history-first-prompts \
--model \
ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--output \
ml/purpose-classifier/outputs/purpose-lite-v1-history-first-prompts-teacher.pt \
--device mps \
--batch-size 16 \
--progress-steps 25 \
--overwrite-output
```
Then run two validation-selected distilled continuation epochs:
```bash
ml/purpose-classifier/venv/bin/python -u \
ml/purpose-classifier/train_deep_mlx.py \
--variant base \
--resume-from \
ml/purpose-classifier/outputs/purpose-deep-v1-base-mlx-cont-3e/model \
--dataset-dir \
ml/purpose-classifier/.artifacts/dataset-v1-history-first-prompts \
--distillation-cache \
ml/purpose-classifier/outputs/purpose-lite-v1-history-first-prompts-teacher.pt \
--distillation-weight 0.5 \
--distillation-temperature 2 \
--epochs 2 \
--learning-rate 1e-5 \
--early-stopping-patience 1 \
--progress-steps 10 \
--output-dir \
ml/purpose-classifier/outputs/purpose-deep-v1-base-mlx-distilled \
--overwrite-output
```
The trainer records the resumed checkpoint as epoch zero before updating anything, so a
distillation regression cannot overwrite the 80.31% candidate. Teacher agreement is
reported for diagnosis but does not enter deep checkpoint selection; overall and hard
primary label accuracy remain the only selection inputs.
Do not launch the large rung yet. It is justified only after base is evaluated on the
frozen set; large must beat base by at least two hard-slice points, while deep itself must
reach 97% scored overall and beat the shipping lite artifact by five hard-slice points.