diff --git a/datagen-prompt-2.md b/datagen-prompt-2.md new file mode 100644 index 0000000..6bfd463 --- /dev/null +++ b/datagen-prompt-2.md @@ -0,0 +1,87 @@ +# Corrective-batch generation prompt (round 2) + +Round 1 (`data/purpose-prompts.jsonl`, 9,217 records) was validated with +`validate-data.py`; this prompt generates a *complementary* set that repairs the measured +deficits. The quotas below are computed so that the MERGED dataset (~12,217 records) lands +on the base prompt's aggregate targets. Because these batches are deliberately +counter-skewed, `validate-data.py` per-batch *share warnings* are expected on them; +per-batch *errors* (openers, length floors, format) must still be zero. Separately, in +curation (not generation): delete the 3 fixture-duplicate lines round 1 shipped, resolve +the 17-record batch remainder, and write the generation manifest. + +--- + +You are generating ROUND 2 of a labeled dataset of prompts that software developers type +into a coding-agent app (like Claude Code) to START a new chat. Follow every rule in the +round-1 brief — first-messages only, the 8 labels and their boundary rules, the JSONL +schema {prompt, purpose, secondary, mixed, difficulty, slice, lang}, the anti-patterns +(no label leakage, no mid-conversation replies, no assistant-directed meta) — EXCEPT +where this brief overrides it. Round 1 drifted in measurable ways; round 2 exists to pull +the merged dataset back on target, so these overrides are hard requirements, not +preferences. + +## What round 1 got wrong (so you don't repeat it) + +1. It opened prompts with the same verbs constantly: "Plan…", "Implement…", "Review…", + "Explain…", "Add…", "Build…", "Audit…". 134 windows broke the opener cap. +2. It wrote almost everything at comfortable middle length — too few terse prompts, far + too few long ones, and its "pasted-context" examples were mostly under 100 tokens. +3. It over-produced backendImpl and core-slice examples; under-produced pasted-context, + mixed, boundary, refactor, review, and writing. +4. It stayed 99% English in most batches. + +## Hard per-batch quotas (each batch = exactly 200 lines) + +Slice counts per batch (deliberately counter-skewed; do not "fix" them toward the +round-1 targets): +- `core`: 58 +- `boundary`: 52 +- `pasted-context`: 50 — and every one's pasted material is genuinely 100–400 tokens + (real-looking stack traces, log tails, failing-test output, diff hunks, ticket text). + Count the tokens; round 1's 60–99-token "pastes" are rejected. +- `mixed`: 24 (`mixed: true`, secondary set — this is the 12% cap exactly) +- `vague-eval`: 16 + +Primary-label counts per batch (non-vague, sums to 184 — rebalances round 1's skew): +- writing: 30, review: 29, refactor: 29, quickFix: 25, frontendImpl: 22, planning: 22, + debugging: 18, backendImpl: 9 + +Length floors per batch (both count toward whatever slice they belong to): +- at least 60 prompts under 8 words +- at least 80 prompts over 60 estimated tokens (the 50 pasted-context entries count; + the other 30+ must come from long conversational core/boundary/mixed prompts — + rambling context-setting, multi-requirement asks, bullet-listed briefs) + +Language per batch: 14–18 lines non-English, spread across es, de, fr, pt, zh, ja +(natural developer code-switching with English tech nouns; put the BCP-47 tag in `lang`). + +## Opening-word discipline (this is where round 1 failed hardest) + +- HARD CAP: no opening word may start more than 2 prompts in any 50 consecutive lines. + The words plan, implement, review, explain, add, build, audit, fix, create, write, + refactor, update are on a watch list — treat each as nearly exhausted before you start. +- Reach the same intents through other doors: start from the noun ("the checkout retry + logic is duplicated in three places…"), the symptom ("payments double-charge when…"), + a question ("why does the exporter…", "is there a cleaner way to…"), the artifact + ("this stack trace keeps showing up:"), a stakeholder ("PM wants a writeup of…"), + lowercase mid-thought fragments ("ok so the settings pane…"), a file path + ("src/sync/reconcile.ts is 900 lines now…"), or pasted material first with the ask at + the end. +- Vary sentence *shape*, not just the first token: imperative, question, complaint, + observation + ask, list of constraints, apology-then-ask, all in the mix. + +## Confusion pairs to emphasize in `boundary` (round 1 under-served these) + +refactor↔quickFix (renames, small restructures), review↔writing (explain vs summarize +vs document), writing↔backendImpl (docs about APIs), review↔debugging (why-questions +about behavior that may or may not be broken), planning↔refactor (restructure strategy +vs restructure execution). Keep planning↔backendImpl light — round 1 already produced +plenty. + +## Process + +15 batches of exactly 200 lines (3,000 total). Before each batch, silently pick 3 fresh +domains + 2 registers + 1 emphasized confusion pair; never repeat a combination. After +each batch, verify against THIS brief's quotas — slice counts, label counts, both length +floors, the 2-per-50 opener cap, non-English count — and fix violations before emitting. +JSONL only; no numbering, no fences, no commentary.