5.0 KiB
Corrective-batch generation prompt (round 2)
Round 1 (data/purpose-prompts.jsonl, 9,217 records) was validated with
validate-data.py; this prompt generates a complementary set that repairs the measured
deficits. The quotas below are computed so that the MERGED dataset (~12,217 records) lands
on the base prompt's aggregate targets. Because these batches are deliberately
counter-skewed, validate-data.py per-batch share warnings are expected on them;
per-batch errors (openers, length floors, format) must still be zero. Separately, in
curation (not generation): delete the 3 fixture-duplicate lines round 1 shipped, resolve
the 17-record batch remainder, and write the generation manifest.
You are generating ROUND 2 of a labeled dataset of prompts that software developers type into a coding-agent app (like Claude Code) to START a new chat. Follow every rule in the round-1 brief — first-messages only, the 8 labels and their boundary rules, the JSONL schema {prompt, purpose, secondary, mixed, difficulty, slice, lang}, the anti-patterns (no label leakage, no mid-conversation replies, no assistant-directed meta) — EXCEPT where this brief overrides it. Round 1 drifted in measurable ways; round 2 exists to pull the merged dataset back on target, so these overrides are hard requirements, not preferences.
What round 1 got wrong (so you don't repeat it)
- It opened prompts with the same verbs constantly: "Plan…", "Implement…", "Review…", "Explain…", "Add…", "Build…", "Audit…". 134 windows broke the opener cap.
- It wrote almost everything at comfortable middle length — too few terse prompts, far too few long ones, and its "pasted-context" examples were mostly under 100 tokens.
- It over-produced backendImpl and core-slice examples; under-produced pasted-context, mixed, boundary, refactor, review, and writing.
- It stayed 99% English in most batches.
Hard per-batch quotas (each batch = exactly 200 lines)
Slice counts per batch (deliberately counter-skewed; do not "fix" them toward the round-1 targets):
core: 58boundary: 52pasted-context: 50 — and every one's pasted material is genuinely 100–400 tokens (real-looking stack traces, log tails, failing-test output, diff hunks, ticket text). Count the tokens; round 1's 60–99-token "pastes" are rejected.mixed: 24 (mixed: true, secondary set — this is the 12% cap exactly)vague-eval: 16
Primary-label counts per batch (non-vague, sums to 184 — rebalances round 1's skew):
- writing: 30, review: 29, refactor: 29, quickFix: 25, frontendImpl: 22, planning: 22, debugging: 18, backendImpl: 9
Length floors per batch (both count toward whatever slice they belong to):
- at least 60 prompts under 8 words
- at least 80 prompts over 60 estimated tokens (the 50 pasted-context entries count; the other 30+ must come from long conversational core/boundary/mixed prompts — rambling context-setting, multi-requirement asks, bullet-listed briefs)
Language per batch: 14–18 lines non-English, spread across es, de, fr, pt, zh, ja
(natural developer code-switching with English tech nouns; put the BCP-47 tag in lang).
Opening-word discipline (this is where round 1 failed hardest)
- HARD CAP: no opening word may start more than 2 prompts in any 50 consecutive lines. The words plan, implement, review, explain, add, build, audit, fix, create, write, refactor, update are on a watch list — treat each as nearly exhausted before you start.
- Reach the same intents through other doors: start from the noun ("the checkout retry logic is duplicated in three places…"), the symptom ("payments double-charge when…"), a question ("why does the exporter…", "is there a cleaner way to…"), the artifact ("this stack trace keeps showing up:"), a stakeholder ("PM wants a writeup of…"), lowercase mid-thought fragments ("ok so the settings pane…"), a file path ("src/sync/reconcile.ts is 900 lines now…"), or pasted material first with the ask at the end.
- Vary sentence shape, not just the first token: imperative, question, complaint, observation + ask, list of constraints, apology-then-ask, all in the mix.
Confusion pairs to emphasize in boundary (round 1 under-served these)
refactor↔quickFix (renames, small restructures), review↔writing (explain vs summarize vs document), writing↔backendImpl (docs about APIs), review↔debugging (why-questions about behavior that may or may not be broken), planning↔refactor (restructure strategy vs restructure execution). Keep planning↔backendImpl light — round 1 already produced plenty.
Process
15 batches of exactly 200 lines (3,000 total). Before each batch, silently pick 3 fresh domains + 2 registers + 1 emphasized confusion pair; never repeat a combination. After each batch, verify against THIS brief's quotas — slice counts, label counts, both length floors, the 2-per-50 opener cap, non-English count — and fix violations before emitting. JSONL only; no numbering, no fences, no commentary.