# Purpose classifier This directory is the reproducible data and training pipeline for `docs/PURPOSE_CLASSIFIER.md`. The current slice covers work item 2 and the first part of work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning, temperature calibration, shared confidence thresholds, and the frozen-set accuracy, recall, hard-slice, calibration, and latency report. ## Data contract The canonical generated sources are listed in `data/generation-manifest.json`. `round2-NN.jsonl` files are retained generation batches and intentionally duplicate `purpose-prompts-round2.jsonl`; they are provenance, not additional training input. `prepare_data.py`: - validates the strict generated-record schema; - removes exact and high-overlap word-trigram duplicates; - fails for review if a high-overlap pair has conflicting labels; - keeps shipped fixtures completely outside source data; - holds every `vague-eval` record out of training; - optionally applies a completed, versioned human-review ledger before splitting; - stratifies by primary purpose, slice, and primary language; and - verifies that the deterministic test partition still matches the versioned `data/frozen-test-v1.jsonl`. The frozen test set is the synthetic JSONL plus the 87 classifiable records in `Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's five `general` records are excluded because `general` is deliberately not a model label. The exact membership and hashes are locked in `data/dataset-v1-manifest.json`. ## Prepare From the repository root: ```bash python3 ml/purpose-classifier/validate-data.py python3 ml/purpose-classifier/prepare_data.py python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py' ``` For a newly generated raw 200-record batch, enable batch-shape checks explicitly with `validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200`. The generated train/validation copies land under `.artifacts/dataset-v1/` and are gitignored. A source, curation, seed, or split-policy change that moves the frozen test set fails closed. After reviewing such a change, intentionally version it with: ```bash python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test ``` ## Train purpose-lite Use a dedicated virtual environment. The base model is pinned to a specific `sentence-transformers/all-MiniLM-L6-v2` commit: a 6-layer, 384-dimensional encoder. The training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export can expose a fixed `1 x 128` runtime shape. Long prompts preserve both ends as `[CLS]` + 63 head tokens + `[SEP]` + 62 tail tokens + `[SEP]`; this keeps the ask when it follows a pasted log or stack trace while retaining enough leading context to interpret it. ```bash python3 -m venv ml/purpose-classifier/.venv ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py ``` The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel accelerator training, install the platform's `torch==2.13.0` build using PyTorch's platform selector, then install `requirements-base.txt`. Training writes a local checkpoint, `calibration.json`, and `metrics.json` under `outputs/purpose-lite-v1/`. It selects checkpoints and fits temperature on label-scorable validation records. When deriving nested HIGH/MEDIUM/LOW cutoffs, every `vague-eval` record counts as an abstention miss even if its synthetic label happens to match. The incoming checkpoint is scored and retained as epoch zero, so a continuation run cannot silently replace it with a regression. Validation early stopping defaults to two epochs without an improvement greater than 0.05 points. Continuation training accepts a local checkpoint. `--boundary-weight` is an opt-in, validation-selected loss weight for the measured weakest slice; it does not add held-out fixtures to training: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \ --model ml/purpose-classifier/outputs/purpose-lite-v1/model \ --epochs 3 --learning-rate 3e-6 --warmup-ratio 0 \ --boundary-weight 2 \ --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune \ --overwrite-output ``` For QAT, `--quantization-aware` replaces the model's linear and embedding forwards with straight-through fake quantization matching the shipping QDQ graph: per-tensor uint8 embeddings, per-channel symmetric int8 linear weights, and per-tensor uint8 activations. Parameter names remain unchanged, so the selected checkpoint reopens as an ordinary Transformers model and uses the same `export.py` path. Keep the incoming checkpoint as epoch zero and select QAT only on validation. Training logs progress every 50 batches by default (`--progress-steps 0` disables it), so a long CPU run remains observable: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \ --model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \ --epochs 2 --learning-rate 1e-6 --warmup-ratio 0 \ --early-stopping-patience 1 --boundary-weight 2 --quantization-aware \ --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-qat1 \ --overwrite-output ``` On dataset v1, that validation-selected run produced a 23,148,500-byte int8 graph at 94.88% frozen accuracy (889/937), 94.46% scored-hard accuracy, and 98.19% scorable PyTorch↔ONNX agreement. It is the current quantized candidate, but remains two correct predictions below the 95% gate. A subsequent validation-selected `5e-7` epoch improved int8 validation accuracy from 93.31% to 93.71% but regressed frozen accuracy to 94.34%; it is rejected. Do not continue optimizer-only QAT sweeps on this split. The next model iteration should incorporate reviewed boundary data and be selected on a revised validation/frozen dataset version. To target only the remaining float→int8 decision drift, cache the float teacher in a separate inference process and use its logits for QAT distillation. Keeping teacher and student models out of the same process avoids doubling peak resident memory: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/cache_teacher.py \ --model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \ --output ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \ --overwrite-output ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \ --model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \ --distillation-cache \ ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \ --distillation-weight 0.9 --distillation-temperature 2 \ --distillation-selection-weight 0.5 --quantization-aware \ --epochs 2 --learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \ --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat \ --overwrite-output ``` The cache binds each logit row to normalized prompt hash plus expected label. Training fails closed if either split changes. Selection combines label accuracy with float-teacher agreement, retains the incoming checkpoint as epoch zero, and logs label/distillation loss separately. A 64-record wiring run exercised cache loading, shuffled row alignment, backpropagation, selection, and ordinary checkpoint reload. The current shared CPU runtime then showed severe post-batch throttling, so no full candidate result is claimed from that canary. ### Native Apple Silicon training with MLX Use the MLX backend when training on Apple Silicon. It implements the same six-layer BERT classifier, fixed head-tail tokenization, export-matched QAT graph, cached-teacher distillation, validation selection, and early stopping with native MLX arrays. Fake quantization is decomposed into Metal-supported round, clip, and straight-through-gradient operations, avoiding PyTorch's unsupported MPS fake-quant operator. Selected weights are written back with the original Hugging Face parameter names, so the existing PyTorch `export.py` and `eval.py` paths remain unchanged. Install the additional pinned dependency into the macOS virtual environment: ```bash ml/purpose-classifier/venv/bin/python -m pip install \ -r ml/purpose-classifier/requirements-mlx.txt ``` Before the first full run on a new MLX or Transformers version, run the fail-closed parity check. It requires exact fake-quant primitives, float-logit parity, matching QAT predictions with bounded backend drift, healthy QAT gradients, and an exact Hugging Face → MLX → Hugging Face weight round trip: ```bash ml/purpose-classifier/venv/bin/python ml/purpose-classifier/verify_mlx.py \ --model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model ``` Then run the distilled QAT candidate natively on Metal: ```bash ml/purpose-classifier/venv/bin/python -u ml/purpose-classifier/train_mlx.py \ --model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \ --distillation-cache \ ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \ --distillation-weight 0.9 --distillation-temperature 2 \ --distillation-selection-weight 0.5 --quantization-aware \ --epochs 2 --early-stopping-patience 1 \ --learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \ --progress-steps 1 \ --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx \ --overwrite-output ``` For a wiring smoke test, use a small deterministic prefix: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \ --epochs 1 --max-train-records 64 --max-validation-records 64 \ --output-dir ml/purpose-classifier/outputs/smoke --overwrite-output ``` ## Evaluate ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py ``` The command returns failure unless label-scorable frozen accuracy is at least 95%, every purpose recall is at least 85%, at least 90% of the deliberately context-free `vague-eval` slice resolves LOW, every misroute stays within one routing cost tier, and measured batch-one p95 is at most 20 ms. Use `--no-gate` only for diagnostic runs. ## Export and score ONNX ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/export.py ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \ --onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \ --report ml/purpose-classifier/outputs/purpose-lite-v1/export/int8-frozen-eval.json ``` `export.py` emits fixed-shape opset-17 fp16 and int8-QDQ graphs, a tokenizer/ normalization contract, golden tokenizations, shared calibration config, graph checks, artifact hashes, and a size report. Its default 256-record quantization calibration sample is deterministic and stratified by purpose, slice, and primary language; the export report records the seed, distribution, and prompt hashes. The int8 graph is the ≤25 MiB shipping candidate; the fp16 graph remains the accelerator-oriented conversion input. When scoring ONNX, add `--compare-pytorch` to measure artifact drift against `--model-dir`. The report then includes overall, label-scorable, and per-slice label agreement plus every correct→incorrect, incorrect→correct, and changed-wrong-label transition: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \ --onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \ --compare-pytorch --no-gate ``` ## Audit curation Run the semantic embedding duplicate audit. It also emits the deterministic, purpose/slice/language-stratified 10% human label-and-difficulty review CSV: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/audit_data.py ``` The semantic pass uses the same commit-pinned MiniLM encoder and fixed 128-token input as `purpose-lite`. Similarity only proposes review candidates; it never edits source data or the frozen split automatically. The 18 word-trigram exclusions and the human-review completion rule are recorded in `data/curation-review-v1.json`; the semantic report is versioned as `data/semantic-audit-v1.json`. ## Optional human review The dataset owner accepted the curated generated labels and difficulty metadata as-is on 2026-07-31, so the blank 1,219-row review sample is not a training or rollout blocker. It remains available as an optional future audit. Check its progress without running the embedding audit again: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/review_data.py ``` Mark each row `accept`, `relabel`, or `reject`. `accept` and `reject` leave the four `reviewed*` fields blank; `reject` requires notes. For `relabel`, blank reviewed fields retain their generated value, `` clears a secondary purpose, and notes are required. If a secondary purpose is added or removed, set `reviewedSlice` consistently (`mixed` when a secondary is present). The validator rejects stale generated columns, missing or duplicate sample rows, invalid label combinations, and partially completed rows. When every row has a human decision, write the versionable ledger: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/review_data.py --finalize ``` Build an isolated candidate split first: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/prepare_data.py \ --human-review ml/purpose-classifier/data/human-review-v1.json \ --output-dir ml/purpose-classifier/.artifacts/reviewed-candidate \ --frozen-test ml/purpose-classifier/.artifacts/reviewed-frozen-candidate.jsonl \ --manifest ml/purpose-classifier/.artifacts/reviewed-manifest-candidate.json \ --refresh-frozen-test ``` Inspect the ledger, decision summary, candidate manifest, and split diff. Only then rerun the same command with the three candidate-path overrides removed to intentionally replace the versioned frozen dataset and manifest. `--regenerate` recreates a blank CSV in the current schema and is only appropriate before review begins. The one-time, hardware-bound energy and accelerator-residency procedure is in `ENERGY_AND_RESIDENCY.md`.