Merge nucleic/upbeat-yarn-seal-ekbe into dev
This commit is contained in:
@@ -29,6 +29,58 @@ The frozen test set is the synthetic JSONL plus the 87 classifiable records in
|
||||
records are excluded because `general` is deliberately not a model label. The exact
|
||||
membership and hashes are locked in `data/dataset-v1-manifest.json`.
|
||||
|
||||
## Full Sol-high label reset
|
||||
|
||||
`rebuild_sol_high.py` is the end-to-end workflow for intentionally invalidating the old
|
||||
labels and rebuilding both classifiers from one teacher. It freezes and relabels every
|
||||
canonical public prompt, every shipped fixture, local Nucleic first-prompt history, and
|
||||
the pinned SWE-chat candidates with `gpt-5.6-sol` at high reasoning effort. The teacher
|
||||
settings are not configurable in this workflow. Batches default to eight prompts and
|
||||
failed structured responses are retried up to ten times.
|
||||
|
||||
The reset is staged below the ignored `.artifacts/sol-high-reset/` directory. Snapshot
|
||||
hashes prevent a resumed run from silently mixing input revisions or teacher settings.
|
||||
Run these commands from the repository root with the purpose-classifier environment:
|
||||
|
||||
```bash
|
||||
"$PY" ml/purpose-classifier/rebuild_sol_high.py snapshot
|
||||
|
||||
"$PY" ml/purpose-classifier/rebuild_sol_high.py label \
|
||||
--python "$PY"
|
||||
|
||||
"$PY" ml/purpose-classifier/rebuild_sol_high.py status
|
||||
```
|
||||
|
||||
Rerunning `label` resumes every existing decision log and does not relabel completed
|
||||
lines. Once `status` reports `"complete": true`, promote the staged result explicitly:
|
||||
|
||||
```bash
|
||||
"$PY" ml/purpose-classifier/rebuild_sol_high.py promote \
|
||||
--confirm overwrite-all-labels-with-sol-high
|
||||
```
|
||||
|
||||
Promotion first validates and curates the complete staged population. It then replaces
|
||||
the two canonical source files, regenerates all round-two mirrors, relabels shipped
|
||||
fixtures (teacher-rejected fixtures become the runtime `general` fallback), rebuilds the
|
||||
frozen public evaluation split, and replaces `.artifacts/dataset-v1/` with the combined
|
||||
training dataset. Nucleic history and SWE-chat are deduplicated against public evaluation
|
||||
data and added to training only; `vague-eval` records remain optimization exclusions.
|
||||
The prior files are moved into a timestamped ignored backup before replacement.
|
||||
|
||||
The old human/semantic review assertions are marked superseded because they were tied to
|
||||
the invalidated label population. Consequently, Sol-high v2 evaluation measures agreement
|
||||
with the new teacher and is not directly comparable with the old v1 scores.
|
||||
|
||||
Print the two clean, from-pretrained-base training commands after promotion:
|
||||
|
||||
```bash
|
||||
"$PY" ml/purpose-classifier/rebuild_sol_high.py train-commands
|
||||
```
|
||||
|
||||
The purpose-lite command starts from the pinned MiniLM revision. The purpose-deep command
|
||||
starts from the pinned ModernBERT base variant; neither command supplies a prior classifier
|
||||
checkpoint or continuation flag.
|
||||
|
||||
## Label local Nucleic history
|
||||
|
||||
`export_nucleic_prompts.py` extracts the first `userText` event from every local
|
||||
|
||||
Reference in New Issue
Block a user