diff --git a/datagen-prompt.md b/datagen-prompt.md index cfcdf0a..fbb56d8 100644 --- a/datagen-prompt.md +++ b/datagen-prompt.md @@ -9,10 +9,15 @@ pasted into the generator's context (contamination). --- You are generating a labeled dataset of prompts that software developers type into a -coding-agent app (like Claude Code) to start or continue a chat. Each example is ONE -prompt a real developer might send, labeled with what the prompt is FOR. The data trains -a small on-device classifier, so realism and diversity matter more than polish; label -precision matters more than anything. +coding-agent app (like Claude Code) to START a new chat. Each example is the OPENING +message of a fresh chat — never a reply inside an ongoing conversation — labeled with +what the prompt is FOR. This matters: the classifier runs exactly once, on the first +message, to pick the chat's model (which then stays fixed for the chat's whole life), so +the training distribution must be first-messages only. A first message may still +reference prior work the way developers really do ("continuing from yesterday's auth +refactor, …", "picking up the payment-flow bug again"), because people routinely start +fresh chats mid-project. The data trains a small on-device classifier, so realism and +diversity matter more than polish; label precision matters more than anything. ## Labels (choose the primary purpose; definitions are exhaustive) @@ -64,9 +69,10 @@ Boundary rules (apply in this order when two labels tempt you): - `pasted-context` — the ask is buried in pasted material: stack traces, log tail, diff hunk, failing test output, a TODO list. Reproduce the pasted material realistically, 100–400 tokens (≈10%) - - `vague-eval` — terse/ambiguous prompts with no recoverable purpose ("continue", - "make it pop", "do the thing we discussed"). Label with your best guess anyway; - these are held out of training by the pipeline (≈5%) + - `vague-eval` — terse/ambiguous prompts with no recoverable purpose, typed as the + first message of a new chat ("continue", "make it pop", "do the thing we + discussed" — real behavior when someone reopens work in a fresh chat). Label with + your best guess anyway; these are held out of training by the pipeline (≈5%) ## Diversity requirements (enforced per batch) @@ -97,6 +103,11 @@ Boundary rules (apply in this order when two labels tempt you): (except the `vague-eval` slice, where that's the point). - Assistant-directed meta ("classify this prompt", "as an AI") — these are prompts TO a coding agent, about code. +- Mid-conversation replies: anything that only makes sense as turn 2+ of a thread — + reacting to an assistant's previous answer ("yes do option 2", "that didn't work, try + again", "same error as before", "looks good, ship it", "no, the OTHER function"). + Referencing prior *work* in a fresh chat is fine (see the intro); referencing a prior + *turn of this conversation* is not, because the classifier never sees those. ## Process