Merge nucleic/eager-glass-wren-mrb5 into dev

This commit is contained in:
2026-07-29 21:01:20 -07:00
parent 728766f1a5
commit a0f34b89cb
+18 -7
View File
@@ -9,10 +9,15 @@ pasted into the generator's context (contamination).
--- ---
You are generating a labeled dataset of prompts that software developers type into a You are generating a labeled dataset of prompts that software developers type into a
coding-agent app (like Claude Code) to start or continue a chat. Each example is ONE coding-agent app (like Claude Code) to START a new chat. Each example is the OPENING
prompt a real developer might send, labeled with what the prompt is FOR. The data trains message of a fresh chat — never a reply inside an ongoing conversation — labeled with
a small on-device classifier, so realism and diversity matter more than polish; label what the prompt is FOR. This matters: the classifier runs exactly once, on the first
precision matters more than anything. message, to pick the chat's model (which then stays fixed for the chat's whole life), so
the training distribution must be first-messages only. A first message may still
reference prior work the way developers really do ("continuing from yesterday's auth
refactor, …", "picking up the payment-flow bug again"), because people routinely start
fresh chats mid-project. The data trains a small on-device classifier, so realism and
diversity matter more than polish; label precision matters more than anything.
## Labels (choose the primary purpose; definitions are exhaustive) ## Labels (choose the primary purpose; definitions are exhaustive)
@@ -64,9 +69,10 @@ Boundary rules (apply in this order when two labels tempt you):
- `pasted-context` — the ask is buried in pasted material: stack traces, log tail, - `pasted-context` — the ask is buried in pasted material: stack traces, log tail,
diff hunk, failing test output, a TODO list. Reproduce the pasted material diff hunk, failing test output, a TODO list. Reproduce the pasted material
realistically, 100–400 tokens (≈10%) realistically, 100–400 tokens (≈10%)
- `vague-eval` — terse/ambiguous prompts with no recoverable purpose ("continue", - `vague-eval` — terse/ambiguous prompts with no recoverable purpose, typed as the
"make it pop", "do the thing we discussed"). Label with your best guess anyway; first message of a new chat ("continue", "make it pop", "do the thing we
these are held out of training by the pipeline (≈5%) discussed" — real behavior when someone reopens work in a fresh chat). Label with
your best guess anyway; these are held out of training by the pipeline (≈5%)
## Diversity requirements (enforced per batch) ## Diversity requirements (enforced per batch)
@@ -97,6 +103,11 @@ Boundary rules (apply in this order when two labels tempt you):
(except the `vague-eval` slice, where that's the point). (except the `vague-eval` slice, where that's the point).
- Assistant-directed meta ("classify this prompt", "as an AI") — these are prompts TO a - Assistant-directed meta ("classify this prompt", "as an AI") — these are prompts TO a
coding agent, about code. coding agent, about code.
- Mid-conversation replies: anything that only makes sense as turn 2+ of a thread —
reacting to an assistant's previous answer ("yes do option 2", "that didn't work, try
again", "same error as before", "looks good, ship it", "no, the OTHER function").
Referencing prior *work* in a fresh chat is fine (see the intro); referencing a prior
*turn of this conversation* is not, because the classifier never sees those.
## Process ## Process