Merge nucleic/eager-glass-wren-mrb5 into dev
This commit is contained in:
@@ -0,0 +1,87 @@
|
||||
# Corrective-batch generation prompt (round 2)
|
||||
|
||||
Round 1 (`data/purpose-prompts.jsonl`, 9,217 records) was validated with
|
||||
`validate-data.py`; this prompt generates a *complementary* set that repairs the measured
|
||||
deficits. The quotas below are computed so that the MERGED dataset (~12,217 records) lands
|
||||
on the base prompt's aggregate targets. Because these batches are deliberately
|
||||
counter-skewed, `validate-data.py` per-batch *share warnings* are expected on them;
|
||||
per-batch *errors* (openers, length floors, format) must still be zero. Separately, in
|
||||
curation (not generation): delete the 3 fixture-duplicate lines round 1 shipped, resolve
|
||||
the 17-record batch remainder, and write the generation manifest.
|
||||
|
||||
---
|
||||
|
||||
You are generating ROUND 2 of a labeled dataset of prompts that software developers type
|
||||
into a coding-agent app (like Claude Code) to START a new chat. Follow every rule in the
|
||||
round-1 brief — first-messages only, the 8 labels and their boundary rules, the JSONL
|
||||
schema {prompt, purpose, secondary, mixed, difficulty, slice, lang}, the anti-patterns
|
||||
(no label leakage, no mid-conversation replies, no assistant-directed meta) — EXCEPT
|
||||
where this brief overrides it. Round 1 drifted in measurable ways; round 2 exists to pull
|
||||
the merged dataset back on target, so these overrides are hard requirements, not
|
||||
preferences.
|
||||
|
||||
## What round 1 got wrong (so you don't repeat it)
|
||||
|
||||
1. It opened prompts with the same verbs constantly: "Plan…", "Implement…", "Review…",
|
||||
"Explain…", "Add…", "Build…", "Audit…". 134 windows broke the opener cap.
|
||||
2. It wrote almost everything at comfortable middle length — too few terse prompts, far
|
||||
too few long ones, and its "pasted-context" examples were mostly under 100 tokens.
|
||||
3. It over-produced backendImpl and core-slice examples; under-produced pasted-context,
|
||||
mixed, boundary, refactor, review, and writing.
|
||||
4. It stayed 99% English in most batches.
|
||||
|
||||
## Hard per-batch quotas (each batch = exactly 200 lines)
|
||||
|
||||
Slice counts per batch (deliberately counter-skewed; do not "fix" them toward the
|
||||
round-1 targets):
|
||||
- `core`: 58
|
||||
- `boundary`: 52
|
||||
- `pasted-context`: 50 — and every one's pasted material is genuinely 100–400 tokens
|
||||
(real-looking stack traces, log tails, failing-test output, diff hunks, ticket text).
|
||||
Count the tokens; round 1's 60–99-token "pastes" are rejected.
|
||||
- `mixed`: 24 (`mixed: true`, secondary set — this is the 12% cap exactly)
|
||||
- `vague-eval`: 16
|
||||
|
||||
Primary-label counts per batch (non-vague, sums to 184 — rebalances round 1's skew):
|
||||
- writing: 30, review: 29, refactor: 29, quickFix: 25, frontendImpl: 22, planning: 22,
|
||||
debugging: 18, backendImpl: 9
|
||||
|
||||
Length floors per batch (both count toward whatever slice they belong to):
|
||||
- at least 60 prompts under 8 words
|
||||
- at least 80 prompts over 60 estimated tokens (the 50 pasted-context entries count;
|
||||
the other 30+ must come from long conversational core/boundary/mixed prompts —
|
||||
rambling context-setting, multi-requirement asks, bullet-listed briefs)
|
||||
|
||||
Language per batch: 14–18 lines non-English, spread across es, de, fr, pt, zh, ja
|
||||
(natural developer code-switching with English tech nouns; put the BCP-47 tag in `lang`).
|
||||
|
||||
## Opening-word discipline (this is where round 1 failed hardest)
|
||||
|
||||
- HARD CAP: no opening word may start more than 2 prompts in any 50 consecutive lines.
|
||||
The words plan, implement, review, explain, add, build, audit, fix, create, write,
|
||||
refactor, update are on a watch list — treat each as nearly exhausted before you start.
|
||||
- Reach the same intents through other doors: start from the noun ("the checkout retry
|
||||
logic is duplicated in three places…"), the symptom ("payments double-charge when…"),
|
||||
a question ("why does the exporter…", "is there a cleaner way to…"), the artifact
|
||||
("this stack trace keeps showing up:"), a stakeholder ("PM wants a writeup of…"),
|
||||
lowercase mid-thought fragments ("ok so the settings pane…"), a file path
|
||||
("src/sync/reconcile.ts is 900 lines now…"), or pasted material first with the ask at
|
||||
the end.
|
||||
- Vary sentence *shape*, not just the first token: imperative, question, complaint,
|
||||
observation + ask, list of constraints, apology-then-ask, all in the mix.
|
||||
|
||||
## Confusion pairs to emphasize in `boundary` (round 1 under-served these)
|
||||
|
||||
refactor↔quickFix (renames, small restructures), review↔writing (explain vs summarize
|
||||
vs document), writing↔backendImpl (docs about APIs), review↔debugging (why-questions
|
||||
about behavior that may or may not be broken), planning↔refactor (restructure strategy
|
||||
vs restructure execution). Keep planning↔backendImpl light — round 1 already produced
|
||||
plenty.
|
||||
|
||||
## Process
|
||||
|
||||
15 batches of exactly 200 lines (3,000 total). Before each batch, silently pick 3 fresh
|
||||
domains + 2 registers + 1 emphasized confusion pair; never repeat a combination. After
|
||||
each batch, verify against THIS brief's quotas — slice counts, label counts, both length
|
||||
floors, the 2-per-50 opener cap, non-English count — and fix violations before emitting.
|
||||
JSONL only; no numbering, no fences, no commentary.
|
||||
Reference in New Issue
Block a user