Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-31 01:24:01 -07:00
parent 0bee416a88
commit 9a1228efbb
7 changed files with 1876 additions and 0 deletions
+47
View File
@@ -283,6 +283,53 @@ the accepted baseline's float checkpoint. That gain does not survive export: the
baseline and fails the 95% shipping gate. Preserve the artifact as a rejected experiment;
`purpose-lite-v1-distilled-qat-mlx-4e` remains the candidate of record.
## Train purpose-deep
The next classifier tier is an MLX-native ModernBERT multi-task model. It keeps the
primary eight-way purpose output and jointly learns secondary purpose, a mixed-intent
flag, and advisory difficulty. Both upstream rungs are immutable: base is ModernBERT
149M at revision `8949b909ec900327062f0ebf497f51aef5e6f0c8`; large is ModernBERT
395M at revision `45bb4654a4d5aaff24dd11d4781fa46d39bf8c13`.
Before the first run on a new MLX/Transformers version, compare the real pinned backbone
against Hugging Face. The check crosses ModernBERT's local-attention window and fails if
pooled-representation drift exceeds `5e-4`:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/verify_deep_mlx.py \
--variant base
```
Start with the base ablation rung and the validated first-prompt history augmentation.
The trainer downloads the pinned checkpoint on first use, fixes every input at 512 tokens
(`255` head + `254` tail + three special tokens for long prompts), and uses gradient
checkpointing by default. Checkpoint selection is half scored overall accuracy and half
scored hard-slice accuracy; auxiliary heads are reported independently and cannot hide a
primary-purpose regression.
```bash
ml/purpose-classifier/venv/bin/python -u \
ml/purpose-classifier/train_deep_mlx.py \
--variant base \
--dataset-dir \
ml/purpose-classifier/.artifacts/dataset-v1-history-first-prompts \
--epochs 3 --early-stopping-patience 1 \
--progress-steps 10 \
--output-dir ml/purpose-classifier/outputs/purpose-deep-v1-base-mlx \
--overwrite-output
```
The defaults use batch size 4 and learning rate `2e-5` for base (2 and `1e-5` for
large). If unified memory is tight, lower `--batch-size` before disabling gradient
checkpointing. `--device cpu` is diagnostic only: a real 512-token backward pass is
expected to be extremely slow there. Each improved epoch atomically rewrites `model/`
and updates `training-state.json`, so progress is visible and an interrupted run retains
the last selected checkpoint.
Do not launch the large rung yet. It is justified only after base is evaluated on the
frozen set; large must beat base by at least two hard-slice points, while deep itself must
reach 97% scored overall and beat the shipping lite artifact by five hard-slice points.
### Convert and validate Core ML
Core ML Tools no longer maintains the legacy ONNX converter, so the Apple artifact is