From 728766f1a5c8b77faa59022568320f7ce8e93572 Mon Sep 17 00:00:00 2001 From: Nucleic Date: Wed, 29 Jul 2026 19:00:34 -0700 Subject: [PATCH] Merge nucleic/eager-glass-wren-mrb5 into dev --- datagen-prompt.md | 109 ++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 109 insertions(+) create mode 100644 datagen-prompt.md diff --git a/datagen-prompt.md b/datagen-prompt.md new file mode 100644 index 0000000..cfcdf0a --- /dev/null +++ b/datagen-prompt.md @@ -0,0 +1,109 @@ +# Synthetic-data generation prompt for the purpose classifier + +The prompt below is fed verbatim to a frontier-model agent to produce training/eval data +per docs/PURPOSE_CLASSIFIER.md §4.1. Record the generating model, date, and batch topics +in the generation manifest alongside the output. The 82 shipped fixtures +(Tests/NucleicCoreTests/Fixtures/purpose-prompts.json) are eval-only and must NOT be +pasted into the generator's context (contamination). + +--- + +You are generating a labeled dataset of prompts that software developers type into a +coding-agent app (like Claude Code) to start or continue a chat. Each example is ONE +prompt a real developer might send, labeled with what the prompt is FOR. The data trains +a small on-device classifier, so realism and diversity matter more than polish; label +precision matters more than anything. + +## Labels (choose the primary purpose; definitions are exhaustive) + +- `planning` — asking for architecture, design docs, RFCs, migration strategy, + roadmaps, breaking work into milestones. The deliverable is a PLAN or DESIGN, not code. +- `backendImpl` — implementing server/API/data/algorithm/system/CLI code: endpoints, + schemas, migrations, queues, caches, auth flows, parsers, background jobs. +- `frontendImpl` — implementing UI: views, components, styling, layout, animation, + themes, screens, visual polish. If the deliverable is something you SEE, it's frontend. +- `quickFix` — a typo, version bump, config tweak, flag flip, one-liner, or a small + contained bugfix the author already understands. Small scope, known change. +- `refactor` — restructuring without behavior change: rename/extract/split/consolidate/ + dedupe/decouple/simplify. The author expects identical behavior after. +- `debugging` — diagnosing a failure the author does NOT yet understand: crashes, stack + traces, regressions, flaky tests, hangs, leaks, wrong output, "why does X happen". +- `review` — reading/judging/explaining EXISTING code or designs: code review, audits, + "what does X do", "is this safe", comparisons, walkthroughs. No code changes requested. +- `writing` — producing prose: docs, READMEs, commit messages, PR descriptions, release + notes, changelogs, summaries, translations, doc comments. + +Boundary rules (apply in this order when two labels tempt you): +1. "Fix" + author already knows the change → `quickFix`. "Fix" + cause unknown / + symptoms described → `debugging`. +2. Rename/restructure "across the codebase" or preserving behavior → `refactor`, even + though a single rename in one file reads as `quickFix`. +3. Docs/comments/prose about code → `writing`, even when the subject is an API or + backend concept ("update the API docs" is `writing`). +4. "Plan/design/architect X" → `planning` even when X is backend or frontend work. + "Plan and implement X" → primary is `planning`, secondary is the implementation label. +5. A pure question about existing behavior → `review` unless something is BROKEN, then + `debugging`. + +## Output format — strict JSONL, one object per line, no commentary + +{"prompt": "...", "purpose": "", "secondary": "