Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 20:56:46 -07:00
parent e05521482f
commit 9e462a79fc
9 changed files with 832 additions and 8 deletions
+53
View File
@@ -196,6 +196,59 @@ routing-tier-drift gates pass. This is the current accuracy-qualified shipping c
latency and energy/residency still require measurement on the target Apple and Windows
accelerator runtimes.
### Convert and validate Core ML
Core ML Tools no longer maintains the legacy ONNX converter, so the Apple artifact is
converted directly from the selected Hugging Face checkpoint. `convert_coreml.py` uses a
fixed-shape export-only BERT forward to avoid dynamic Transformers masking helpers, checks
that forward against Transformers before conversion, writes an ML Program package, and
records hashes for every package file.
Install the pinned converter in the macOS environment and create the package:
```bash
ml/purpose-classifier/venv/bin/python -m pip install \
-r ml/purpose-classifier/requirements-coreml.txt
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/convert_coreml.py \
--model-dir \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/model \
--output \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
--overwrite-output
```
Run the frozen gate with CPU+Neural Engine placement and compare labels directly with the
accepted int8 ONNX artifact. Gated Core ML evaluation fails closed without
`--compare-onnx`, and requires at least 99.5% scorable label agreement:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/eval.py \
--model-dir \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/model \
--calibration \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/calibration.json \
--coreml-model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
--coreml-compute-units cpu-and-ne \
--compare-onnx \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/export/purpose-lite-v1-int8-qdq.onnx \
--report \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/ane-frozen-eval.json
```
Record the compute plan separately; this reports both operation-count and estimated-cost
ANE shares. Repeat evaluation with `--coreml-compute-units cpu-only --no-gate` before the
energy comparison in `ENERGY_AND_RESIDENCY.md`:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/inspect_coreml.py \
--model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
--compute-units cpu-and-ne \
--report \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/ane-compute-plan.json
```
For a wiring smoke test, use a small deterministic prefix:
```bash