# Purpose classifier This directory is the reproducible data and training pipeline for `docs/PURPOSE_CLASSIFIER.md`. The current slice covers work item 2 and the first part of work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning, temperature calibration, shared confidence thresholds, and the frozen-set accuracy, recall, hard-slice, calibration, and latency report. ## Data contract The canonical generated sources are listed in `data/generation-manifest.json`. `round2-NN.jsonl` files are retained generation batches and intentionally duplicate `purpose-prompts-round2.jsonl`; they are provenance, not additional training input. `prepare_data.py`: - validates the strict generated-record schema; - removes exact and high-overlap word-trigram duplicates; - fails for review if a high-overlap pair has conflicting labels; - keeps shipped fixtures completely outside source data; - holds every `vague-eval` record out of training; - stratifies by primary purpose, slice, and primary language; and - verifies that the deterministic test partition still matches the versioned `data/frozen-test-v1.jsonl`. The frozen test set is the synthetic JSONL plus the 87 classifiable records in `Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's five `general` records are excluded because `general` is deliberately not a model label. The exact membership and hashes are locked in `data/dataset-v1-manifest.json`. ## Prepare From the repository root: ```bash python3 ml/purpose-classifier/validate-data.py python3 ml/purpose-classifier/prepare_data.py python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py' ``` For a newly generated raw 200-record batch, enable batch-shape checks explicitly with `validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200`. The generated train/validation copies land under `.artifacts/dataset-v1/` and are gitignored. A source, curation, seed, or split-policy change that moves the frozen test set fails closed. After reviewing such a change, intentionally version it with: ```bash python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test ``` ## Train purpose-lite Use a dedicated virtual environment. The base model is pinned to a specific `sentence-transformers/all-MiniLM-L6-v2` commit: a 6-layer, 384-dimensional encoder. The training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export can expose a fixed `1 x 128` runtime shape. ```bash python3 -m venv ml/purpose-classifier/.venv ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py ``` The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel accelerator training, install the platform's `torch==2.13.0` build using PyTorch's platform selector, then install `requirements-base.txt`. Training writes a local checkpoint, `calibration.json`, and `metrics.json` under `outputs/purpose-lite-v1/`. It fits one validation-only temperature and derives nested HIGH/MEDIUM/LOW cutoffs from calibrated top-one probability plus top-two margin. For a wiring smoke test, use a small deterministic prefix: ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \ --epochs 1 --max-train-records 64 --max-validation-records 64 \ --output-dir ml/purpose-classifier/outputs/smoke --overwrite-output ``` ## Evaluate ```bash ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py ``` The command returns failure unless frozen accuracy is at least 95%, every purpose recall is at least 85%, and measured batch-one p95 is at most 20 ms. Use `--no-gate` only for diagnostic runs. Accelerator residency, ONNX export/quantization, tokenizer golden tests, tier-drift evaluation, and Core ML parity remain follow-on work. Before calling dataset work item 2 complete, also run a semantic embedding duplicate audit and record the planned 10% human label spot-check; the current dependency-free word-trigram pass is deliberately conservative.