Merge nucleic/sleek-ember-seal-uady into dev
This commit is contained in:
@@ -227,6 +227,53 @@ routing-tier-drift gates pass. This is the current accuracy-qualified shipping c
|
||||
latency and energy/residency still require measurement on the target Apple and Windows
|
||||
accelerator runtimes.
|
||||
|
||||
### First-prompt history augmentation experiment
|
||||
|
||||
`prepare_history_experiment.py` appends the labeled Nucleic first-prompt corpus to
|
||||
training only. It preserves validation and test byte-for-byte, excludes `vague-eval`
|
||||
records from optimization, removes exact base/evaluation overlap, and applies the
|
||||
canonical 0.92 near-duplicate guard against evaluation fixtures and earlier history
|
||||
records. Every exclusion is represented only by hashes and source line in
|
||||
`history-exclusions.jsonl`; `manifest.json` binds all input and output hashes.
|
||||
|
||||
Build the augmented split and its teacher cache:
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/venv/bin/python \
|
||||
ml/purpose-classifier/prepare_history_experiment.py
|
||||
ml/purpose-classifier/venv/bin/python -u \
|
||||
ml/purpose-classifier/cache_teacher.py \
|
||||
--dataset-dir \
|
||||
ml/purpose-classifier/.artifacts/dataset-v1-history-first-prompts \
|
||||
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
|
||||
--output \
|
||||
ml/purpose-classifier/outputs/purpose-lite-v1-history-first-prompts-teacher.pt \
|
||||
--device mps --batch-size 16 --progress-steps 25
|
||||
```
|
||||
|
||||
Then run the same validation-selected MLX recipe as the accepted baseline:
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/venv/bin/python -u ml/purpose-classifier/train_mlx.py \
|
||||
--device metal \
|
||||
--dataset-dir \
|
||||
ml/purpose-classifier/.artifacts/dataset-v1-history-first-prompts \
|
||||
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
|
||||
--distillation-cache \
|
||||
ml/purpose-classifier/outputs/purpose-lite-v1-history-first-prompts-teacher.pt \
|
||||
--distillation-weight 0.9 --distillation-temperature 2 \
|
||||
--distillation-selection-weight 0.5 --quantization-aware \
|
||||
--epochs 4 --early-stopping-patience 1 \
|
||||
--learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \
|
||||
--progress-steps 1 \
|
||||
--output-dir \
|
||||
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-history-v1
|
||||
```
|
||||
|
||||
MLX remains Metal-first. `--device cpu` is an explicit diagnostic fallback for parity
|
||||
checks and bounded smoke tests; it is not an acceptable full-training path when Metal is
|
||||
available.
|
||||
|
||||
### Convert and validate Core ML
|
||||
|
||||
Core ML Tools no longer maintains the legacy ONNX converter, so the Apple artifact is
|
||||
|
||||
Reference in New Issue
Block a user