Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 19:21:08 -07:00
parent bb6d53a520
commit 09f98cdd00
7 changed files with 511 additions and 25 deletions
+33 -3
View File
@@ -117,6 +117,34 @@ it is rejected. Do not continue optimizer-only QAT sweeps on this split. The nex
iteration should incorporate reviewed boundary data and be selected on a revised
validation/frozen dataset version.
To target only the remaining float→int8 decision drift, cache the float teacher in a
separate inference process and use its logits for QAT distillation. Keeping teacher and
student models out of the same process avoids doubling peak resident memory:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/cache_teacher.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--output ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \
--overwrite-output
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--distillation-cache \
ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \
--distillation-weight 0.9 --distillation-temperature 2 \
--distillation-selection-weight 0.5 --quantization-aware \
--epochs 2 --learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat \
--overwrite-output
```
The cache binds each logit row to normalized prompt hash plus expected label. Training
fails closed if either split changes. Selection combines label accuracy with float-teacher
agreement, retains the incoming checkpoint as epoch zero, and logs label/distillation loss
separately. A 64-record wiring run exercised cache loading, shuffled row alignment,
backpropagation, selection, and ordinary checkpoint reload. The current shared CPU runtime
then showed severe post-batch throttling, so no full candidate result is claimed from that
canary.
For a wiring smoke test, use a small deterministic prefix:
```bash
@@ -178,10 +206,12 @@ the frozen split automatically. The 18 word-trigram exclusions and the human-rev
completion rule are recorded in `data/curation-review-v1.json`; the semantic report is
versioned as `data/semantic-audit-v1.json`.
## Complete the human review
## Optional human review
The deterministic CSV currently contains 1,219 blank review rows. Check progress without
running the embedding audit again:
The dataset owner accepted the curated generated labels and difficulty metadata as-is on
2026-07-31, so the blank 1,219-row review sample is not a training or rollout blocker. It
remains available as an optional future audit. Check its progress without running the
embedding audit again:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/review_data.py