Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 18:39:33 -07:00
parent e4403b570d
commit ddbff97191
3 changed files with 191 additions and 0 deletions
+26
View File
@@ -90,6 +90,32 @@ ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--overwrite-output
```
For QAT, `--quantization-aware` replaces the model's linear and embedding forwards with
straight-through fake quantization matching the shipping QDQ graph: per-tensor uint8
embeddings, per-channel symmetric int8 linear weights, and per-tensor uint8 activations.
Parameter names remain unchanged, so the selected checkpoint reopens as an ordinary
Transformers model and uses the same `export.py` path. Keep the incoming checkpoint as
epoch zero and select QAT only on validation. Training logs progress every 50 batches by
default (`--progress-steps 0` disables it), so a long CPU run remains observable:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--epochs 2 --learning-rate 1e-6 --warmup-ratio 0 \
--early-stopping-patience 1 --boundary-weight 2 --quantization-aware \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-qat1 \
--overwrite-output
```
On dataset v1, that validation-selected run produced a 23,148,500-byte int8 graph at
94.88% frozen accuracy (889/937), 94.46% scored-hard accuracy, and 98.19% scorable
PyTorch↔ONNX agreement. It is the current quantized candidate, but remains two correct
predictions below the 95% gate. A subsequent validation-selected `5e-7` epoch improved
int8 validation accuracy from 93.31% to 93.71% but regressed frozen accuracy to 94.34%;
it is rejected. Do not continue optimizer-only QAT sweeps on this split. The next model
iteration should incorporate reviewed boundary data and be selected on a revised
validation/frozen dataset version.
For a wiring smoke test, use a small deterministic prefix:
```bash