Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 21:08:05 -07:00
parent 9e462a79fc
commit 40176d0c76
4 changed files with 373 additions and 3 deletions
+24 -2
View File
@@ -217,6 +217,28 @@ ml/purpose-classifier/venv/bin/python ml/purpose-classifier/convert_coreml.py \
--overwrite-output
```
The direct float16 package is the conversion baseline, not the accepted Apple artifact.
On the first physical Apple-Silicon run it scored 94.98% (890/937), two correct decisions
behind the accepted ONNX graph, with 97.97% scorable label agreement. Calibrate a
Core ML-native W8A8 candidate with the same deterministic 256-record sample and QDQ policy
as the ONNX exporter:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/quantize_coreml.py \
--model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
--model-dir \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/model \
--output \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-w8a8.mlpackage \
--overwrite-output
```
Activation calibration is grouped to keep temporary Core ML packages bounded and prints
progress while it runs. The candidate uses per-tensor asymmetric uint8 activations,
per-channel symmetric int8 linear weights, and per-tensor asymmetric uint8 embedding
weights. It fails the command if the resulting package exceeds 25 MiB.
Run the frozen gate with CPU+Neural Engine placement and compare labels directly with the
accepted int8 ONNX artifact. Gated Core ML evaluation fails closed without
`--compare-onnx`, and requires at least 99.5% scorable label agreement:
@@ -228,7 +250,7 @@ ml/purpose-classifier/venv/bin/python ml/purpose-classifier/eval.py \
--calibration \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/calibration.json \
--coreml-model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-w8a8.mlpackage \
--coreml-compute-units cpu-and-ne \
--compare-onnx \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/export/purpose-lite-v1-int8-qdq.onnx \
@@ -243,7 +265,7 @@ energy comparison in `ENERGY_AND_RESIDENCY.md`:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/inspect_coreml.py \
--model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-w8a8.mlpackage \
--compute-units cpu-and-ne \
--report \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/ane-compute-plan.json