Merge nucleic/fuzzy-dewy-urchin-hpvd into dev

This commit is contained in:
2026-08-02 22:04:24 -07:00
parent 0f639bfa05
commit ae616c6f85
2 changed files with 691 additions and 15 deletions
+56 -15
View File
@@ -1,10 +1,10 @@
# Purpose classifier
This directory is the reproducible data and training pipeline for
`docs/PURPOSE_CLASSIFIER.md`. The current slice covers work item 2 and the first part of
work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning,
temperature calibration, shared confidence thresholds, and the frozen-set accuracy,
recall, hard-slice, calibration, and latency report.
`docs/PURPOSE_CLASSIFIER.md`. The current dataset is the promoted Sol-high v2 reset. Its
dataset lineage, completed lite/deep experiments, artifact hashes, decisions, and bounded
next steps are locked in `experiments/sol-high-v2.json`. Older v1 results below are kept
as implementation history and are not label-compatible candidates for Sol-high v2.
## Data contract
@@ -24,8 +24,8 @@ The canonical generated sources are listed in `data/generation-manifest.json`.
- verifies that the deterministic test partition still matches the versioned
`data/frozen-test-v1.jsonl`.
The frozen test set is the synthetic JSONL plus the 87 classifiable records in
`Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's five `general`
The Sol-high v2 frozen test set is the synthetic JSONL plus the 89 classifiable records in
`Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's three `general`
records are excluded because `general` is deliberately not a model label. The exact
membership and hashes are locked in `data/dataset-v1-manifest.json`.
@@ -100,6 +100,41 @@ The purpose-lite command starts from the pinned MiniLM revision. The purpose-dee
starts from the pinned ModernBERT base variant; neither command supplies a prior classifier
checkpoint or continuation flag.
### Sol-high v2 promotion and training result
Promotion completed on 2026-08-02/03 with 12,164 combined training records, 1,208
validation records, and 1,119 synthetic test records. It explicitly excluded the reviewed
real-population conflicts at labeled source lines 1201 and 1481. The ignored combined
training manifest is bound by SHA-256 in the versioned experiment ledger without exposing
private prompt text.
| Run | Selected validation | Frozen evaluation | Decision |
|---|---:|---:|---|
| lite, from pretrained | 91.28% primary | 93.29% overall; 88.84% hard | rejected |
| lite, boundary continuation | 92.04% primary | 93.72% overall; 91.16% hard | rejected; 13 correct decisions short of 95% |
| deep base | 73.39% primary; 67.50% hard | not opened | rejected on validation |
| deep distilled continuation | 77.64% primary; 72.92% hard | not opened | rejected on validation |
The boundary continuation is useful as the selected lite validation baseline and cached
teacher, but it is not a shipping candidate. Do not tune another lite continuation against
this frozen set, start QAT/export, chain another deep continuation, or start ModernBERT
large/purpose-max.
The next result-bearing work is a validation-only representation probe: compare the
current pretrained prediction-head CLS representation with attention-masked mean pooling
using a regularized primary-only linear head. A pooling change earns one full
ModernBERT-base run only if it gains at least three points on validation overall or hard,
reaches at least 85% overall, and causes no per-purpose collapse. Before any deep frozen
look, the revised model must come within two points of lite overall and beat lite by at
least three points on the identical validation hard slice. If the probe fails, compare a
sentence-trained encoder with the same protocol or remove the deep tier; increasing model
size is not the next variable.
For lite, use validation-only error analysis and data improvements, then reserve a new
independent holdout before the next candidate cycle. QAT, export, target-runtime parity,
latency, energy, and residency resume only after a new float candidate qualifies. See
`experiments/sol-high-v2.json` for exact metrics, hashes, and gates.
## Label local Nucleic history
`export_nucleic_prompts.py` extracts the first `userText` event from every local
@@ -333,9 +368,10 @@ The full Metal run stopped after epoch three and selected epoch two at 94.75% fa
validation accuracy. Its 23,148,500-byte int8-QDQ export scores **95.20% frozen
(892/937)**, 95.19% macro recall, 94.17% scored-hard accuracy, and 98.08% scorable
PyTorch↔ONNX agreement. Every purpose recall is above 91%, and the vague-abstention and
routing-tier-drift gates pass. This is the current accuracy-qualified shipping candidate;
latency and energy/residency still require measurement on the target Apple and Windows
accelerator runtimes.
routing-tier-drift gates pass. This became the accepted artifact for the historical v1
label contract; it is not a Sol-high v2 candidate. Its remaining latency and
energy/residency work is relevant only as runtime evidence unless a new candidate adopts
the same export path.
### First-prompt history augmentation experiment
@@ -390,10 +426,15 @@ float checkpoint improves frozen v1 to **95.09% (891/937)**, two correct decisio
the accepted baseline's float checkpoint. That gain does not survive export: the matched
23,148,500-byte int8-QDQ graph scores **94.66% (887/937)**, 94.66% macro recall, and
94.17% scored-hard accuracy. It is five correct decisions behind the accepted int8
baseline and fails the 95% shipping gate. Preserve the artifact as a rejected experiment;
`purpose-lite-v1-distilled-qat-mlx-4e` remains the candidate of record.
baseline and fails the 95% shipping gate. Preserve the artifact as a rejected historical
experiment; `purpose-lite-v1-distilled-qat-mlx-4e` remains the accepted v1 artifact only.
## Train purpose-deep
## Purpose-deep implementation and historical v1 experiments
The Sol-high v2 base and distilled runs described above supersede the experimental
sequence in this section. Keep the commands and results below for implementation lineage;
do not run another continuation or the large rung from them. The representation probe in
`experiments/sol-high-v2.json` is the current deep next step.
The next classifier tier is an MLX-native ModernBERT multi-task model. It keeps the
primary eight-way purpose output and jointly learns secondary purpose, a mixed-intent
@@ -513,9 +554,9 @@ distillation regression cannot overwrite the 80.31% candidate. Teacher agreement
reported for diagnosis but does not enter deep checkpoint selection; overall and hard
primary label accuracy remain the only selection inputs.
Do not launch the large rung yet. It is justified only after base is evaluated on the
frozen set; large must beat base by at least two hard-slice points, while deep itself must
reach 97% scored overall and beat the shipping lite artifact by five hard-slice points.
The historical large-rung rule required base to qualify first, large to beat base by at
least two hard-slice points, and deep to reach 97% scored overall while beating lite by
five hard-slice points. Neither the v1 nor Sol-high v2 base result unlocked that rung.
### Convert and validate Core ML