Merge nucleic/eager-glass-wren-mrb5 into dev
This commit is contained in:
@@ -0,0 +1,109 @@
|
|||||||
|
# Synthetic-data generation prompt for the purpose classifier
|
||||||
|
|
||||||
|
The prompt below is fed verbatim to a frontier-model agent to produce training/eval data
|
||||||
|
per docs/PURPOSE_CLASSIFIER.md §4.1. Record the generating model, date, and batch topics
|
||||||
|
in the generation manifest alongside the output. The 82 shipped fixtures
|
||||||
|
(Tests/NucleicCoreTests/Fixtures/purpose-prompts.json) are eval-only and must NOT be
|
||||||
|
pasted into the generator's context (contamination).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
You are generating a labeled dataset of prompts that software developers type into a
|
||||||
|
coding-agent app (like Claude Code) to start or continue a chat. Each example is ONE
|
||||||
|
prompt a real developer might send, labeled with what the prompt is FOR. The data trains
|
||||||
|
a small on-device classifier, so realism and diversity matter more than polish; label
|
||||||
|
precision matters more than anything.
|
||||||
|
|
||||||
|
## Labels (choose the primary purpose; definitions are exhaustive)
|
||||||
|
|
||||||
|
- `planning` — asking for architecture, design docs, RFCs, migration strategy,
|
||||||
|
roadmaps, breaking work into milestones. The deliverable is a PLAN or DESIGN, not code.
|
||||||
|
- `backendImpl` — implementing server/API/data/algorithm/system/CLI code: endpoints,
|
||||||
|
schemas, migrations, queues, caches, auth flows, parsers, background jobs.
|
||||||
|
- `frontendImpl` — implementing UI: views, components, styling, layout, animation,
|
||||||
|
themes, screens, visual polish. If the deliverable is something you SEE, it's frontend.
|
||||||
|
- `quickFix` — a typo, version bump, config tweak, flag flip, one-liner, or a small
|
||||||
|
contained bugfix the author already understands. Small scope, known change.
|
||||||
|
- `refactor` — restructuring without behavior change: rename/extract/split/consolidate/
|
||||||
|
dedupe/decouple/simplify. The author expects identical behavior after.
|
||||||
|
- `debugging` — diagnosing a failure the author does NOT yet understand: crashes, stack
|
||||||
|
traces, regressions, flaky tests, hangs, leaks, wrong output, "why does X happen".
|
||||||
|
- `review` — reading/judging/explaining EXISTING code or designs: code review, audits,
|
||||||
|
"what does X do", "is this safe", comparisons, walkthroughs. No code changes requested.
|
||||||
|
- `writing` — producing prose: docs, READMEs, commit messages, PR descriptions, release
|
||||||
|
notes, changelogs, summaries, translations, doc comments.
|
||||||
|
|
||||||
|
Boundary rules (apply in this order when two labels tempt you):
|
||||||
|
1. "Fix" + author already knows the change → `quickFix`. "Fix" + cause unknown /
|
||||||
|
symptoms described → `debugging`.
|
||||||
|
2. Rename/restructure "across the codebase" or preserving behavior → `refactor`, even
|
||||||
|
though a single rename in one file reads as `quickFix`.
|
||||||
|
3. Docs/comments/prose about code → `writing`, even when the subject is an API or
|
||||||
|
backend concept ("update the API docs" is `writing`).
|
||||||
|
4. "Plan/design/architect X" → `planning` even when X is backend or frontend work.
|
||||||
|
"Plan and implement X" → primary is `planning`, secondary is the implementation label.
|
||||||
|
5. A pure question about existing behavior → `review` unless something is BROKEN, then
|
||||||
|
`debugging`.
|
||||||
|
|
||||||
|
## Output format — strict JSONL, one object per line, no commentary
|
||||||
|
|
||||||
|
{"prompt": "...", "purpose": "<primary label>", "secondary": "<label or null>",
|
||||||
|
"mixed": <bool>, "difficulty": <0.0-1.0>, "slice": "<slice tag>", "lang": "<bcp47>"}
|
||||||
|
|
||||||
|
- `secondary`/`mixed`: only when the prompt genuinely asks for two purposes ("plan and
|
||||||
|
implement…", "fix the crash and write a regression test note"). At most ~12% of
|
||||||
|
examples; `mixed` false → `secondary` null.
|
||||||
|
- `difficulty`: how much model capability the TASK described would need. Rubric:
|
||||||
|
0.0–0.2 trivial (typo, one-liner); 0.3–0.5 routine scoped work; 0.6–0.8 multi-file /
|
||||||
|
multi-constraint / gnarly diagnosis; 0.9–1.0 long-horizon, architectural, high-risk.
|
||||||
|
Grade the task, not the prompt's length.
|
||||||
|
- `slice`, one of:
|
||||||
|
- `core` — clear, single-purpose prompts (≈55% of output)
|
||||||
|
- `boundary` — deliberately near a label boundary per the rules above (≈20%)
|
||||||
|
- `mixed` — genuine two-purpose prompts (≈10%)
|
||||||
|
- `pasted-context` — the ask is buried in pasted material: stack traces, log tail,
|
||||||
|
diff hunk, failing test output, a TODO list. Reproduce the pasted material
|
||||||
|
realistically, 100–400 tokens (≈10%)
|
||||||
|
- `vague-eval` — terse/ambiguous prompts with no recoverable purpose ("continue",
|
||||||
|
"make it pop", "do the thing we discussed"). Label with your best guess anyway;
|
||||||
|
these are held out of training by the pipeline (≈5%)
|
||||||
|
|
||||||
|
## Diversity requirements (enforced per batch)
|
||||||
|
|
||||||
|
- Length: from 2 words to ~400 tokens; at least 15% under 8 words, at least 15% over
|
||||||
|
60 tokens.
|
||||||
|
- Register: terse imperatives, polite asks, stream-of-consciousness, bullet lists,
|
||||||
|
mid-thought fragments, sloppy typing with real typos (don't correct them), pasted
|
||||||
|
Slack/ticket text. NEVER the same opening verb more than 3 times per 50 examples.
|
||||||
|
- Domain: rotate web frontend, iOS/macOS (SwiftUI), Android, backend (Go/Rust/Python/
|
||||||
|
TS/Java), data/ML, infra/DevOps, embedded, games, databases, CLI tools.
|
||||||
|
- Tech nouns: use real, current, varied technology names and file paths; invent
|
||||||
|
plausible project-specific names (components, services, feature flags) so examples
|
||||||
|
aren't keyword-matchable.
|
||||||
|
- Language: ~95% English; ~5% spread across es, de, fr, pt, zh, ja (natural developer
|
||||||
|
usage, often code-switched with English tech nouns).
|
||||||
|
- Confusion pairs to mine deliberately in the `boundary` slice, several dozen each:
|
||||||
|
writing↔backendImpl (docs about APIs), quickFix↔refactor (renames), review↔debugging
|
||||||
|
(why-questions), planning↔backendImpl ("plan and implement"), frontendImpl↔quickFix
|
||||||
|
(small UI tweaks), review↔writing (explain vs summarize).
|
||||||
|
|
||||||
|
## Anti-patterns (rejected in QA)
|
||||||
|
|
||||||
|
- Label leakage: prompts must never contain the label word used AS a label hint
|
||||||
|
("refactor this" is fine and common; "this is a refactor task:" is not).
|
||||||
|
- Template smell: recycled sentence skeletons with one noun swapped; enumerated
|
||||||
|
"Task 47:" prefixes; uniform lengths; every example ending in a period.
|
||||||
|
- Impossible labels: prompts a human labeler couldn't defend from the text alone
|
||||||
|
(except the `vague-eval` slice, where that's the point).
|
||||||
|
- Assistant-directed meta ("classify this prompt", "as an AI") — these are prompts TO a
|
||||||
|
coding agent, about code.
|
||||||
|
|
||||||
|
## Process
|
||||||
|
|
||||||
|
Produce the dataset in batches of 200 lines. Before each batch, silently pick a fresh
|
||||||
|
combination of 3 domains + 2 registers + 1 confusion pair to emphasize, so no two
|
||||||
|
batches have the same texture (do not print your picks — JSONL lines only). After each
|
||||||
|
batch, self-check against the anti-patterns and the slice/label distributions, and fix
|
||||||
|
violations before emitting. Across the full run, keep primary labels within ±15% of
|
||||||
|
uniform across the 8 classes (the vague-eval slice is exempt). Target total: 8,000
|
||||||
|
lines. Do not number examples. Do not wrap output in markdown fences. JSONL only.
|
||||||
Reference in New Issue
Block a user