Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 20:42:39 -07:00
parent 90092332db
commit e05521482f
+11 -3
View File
@@ -110,7 +110,7 @@ ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
On dataset v1, that validation-selected run produced a 23,148,500-byte int8 graph at On dataset v1, that validation-selected run produced a 23,148,500-byte int8 graph at
94.88% frozen accuracy (889/937), 94.46% scored-hard accuracy, and 98.19% scorable 94.88% frozen accuracy (889/937), 94.46% scored-hard accuracy, and 98.19% scorable
PyTorch↔ONNX agreement. It is the current quantized candidate, but remains two correct PyTorch↔ONNX agreement. It was the pre-distillation quantized candidate and remained two correct
predictions below the 95% gate. A subsequent validation-selected `5e-7` epoch improved predictions below the 95% gate. A subsequent validation-selected `5e-7` epoch improved
int8 validation accuracy from 93.31% to 93.71% but regressed frozen accuracy to 94.34%; int8 validation accuracy from 93.31% to 93.71% but regressed frozen accuracy to 94.34%;
it is rejected. Do not continue optimizer-only QAT sweeps on this split. The next model it is rejected. Do not continue optimizer-only QAT sweeps on this split. The next model
@@ -181,13 +181,21 @@ ml/purpose-classifier/venv/bin/python -u ml/purpose-classifier/train_mlx.py \
ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \ ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \
--distillation-weight 0.9 --distillation-temperature 2 \ --distillation-weight 0.9 --distillation-temperature 2 \
--distillation-selection-weight 0.5 --quantization-aware \ --distillation-selection-weight 0.5 --quantization-aware \
--epochs 2 --early-stopping-patience 1 \ --epochs 4 --early-stopping-patience 1 \
--learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \ --learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \
--progress-steps 1 \ --progress-steps 1 \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx \ --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e \
--overwrite-output --overwrite-output
``` ```
The full Metal run stopped after epoch three and selected epoch two at 94.75% fake-quant
validation accuracy. Its 23,148,500-byte int8-QDQ export scores **95.20% frozen
(892/937)**, 95.19% macro recall, 94.17% scored-hard accuracy, and 98.08% scorable
PyTorch↔ONNX agreement. Every purpose recall is above 91%, and the vague-abstention and
routing-tier-drift gates pass. This is the current accuracy-qualified shipping candidate;
latency and energy/residency still require measurement on the target Apple and Windows
accelerator runtimes.
For a wiring smoke test, use a small deterministic prefix: For a wiring smoke test, use a small deterministic prefix:
```bash ```bash