Files
nucleic-purpose-classifier/data/opus-30.jsonl
T
2026-07-29 23:45:26 -07:00

132 lines
30 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{"prompt": "sms gateway: multi-carrier routing with per-carrier throughput limits, delivery receipts reconciled back to the message, and a failover when a carrier starts erroring", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "the messaging gateway. three carriers, each with a different api and its own idea of what a delivery receipt means, and our routing is a hardcoded if-chain on country code. i want the plan for proper routing with cost and quality inputs, plus how we handle a carrier being down at 3am without a human", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "message status timeline in the console: accepted, submitted to carrier, delivered or failed, with the carrier's own status shown alongside ours", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "the per-carrier rate limit is a single global counter", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "core", "lang": "en"}
{"prompt": "our three carrier adapters each map status codes to our internal states differently, so 'delivered' means three things. one mapping table, and report which historical messages would be classified differently", "purpose": "refactor", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "about 2% of messages sit in 'submitted' forever with no receipt ever arriving", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "explain how our retry interacts with carrier-side deduplication — can a retried message be delivered twice", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "messaging api documentation: the status lifecycle, what each status guarantees, the retry behavior, and the receipt webhook", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "boundary", "lang": "en"}
{"prompt": "delivery receipt reconciliation job that closes out messages with no receipt after the carrier's stated window, marking them unknown rather than delivered", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "the reconciler marks unknown messages as delivered so our delivery rate looks great and is a lie", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "boundary", "lang": "en"}
{"prompt": "supervision tree for the carrier connections so one carrier's adapter crashing doesn't take the others down", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "one carrier adapter crashing takes the whole gateway down for 20 seconds", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "review our genserver state — i think the message queue is held in process memory and lost on a crash", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "figma plugin: read the selected frames and export our design tokens as json in the format our build expects", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "the plugin ui needs a token list with a diff against what's currently in the repo, and an export button per group", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "the plugin exports color values as rgba floats and our build wants hex", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "core", "lang": "en"}
{"prompt": "the plugin hangs on files with more than about 2000 layers", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "plan the design-to-code handoff properly — tokens, component naming, and how a designer knows what's already built", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "our token naming in figma and in code differ, so the export needs a mapping file that's maintained by hand. align the names in both, no visual change", "purpose": "refactor", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "boundary", "lang": "en"}
{"prompt": "explain how the plugin authenticates to our repo, i want to know if a designer's token is being stored in the plugin", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "plugin usage docs for designers: what it exports, what it can't, and the naming rules it depends on", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.5, "slice": "boundary", "lang": "en"}
{"prompt": "postgres ha: patroni with three nodes, synchronous replication to one, and the failover tested rather than assumed", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "we have one postgres, nightly backups to s3, and no replica. that's the whole story. i want the plan for something we can stand behind, with the rpo and rto per option and what each costs", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "the backup job hasn't succeeded in 11 days and nothing alerted", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "core", "lang": "en"}
{"prompt": "ok next", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "vague-eval", "lang": "en"}
{"prompt": "here's the failover test result from last night, and it did not go well:\n\n21:04 killed primary (db-01) with kill -9 on postgres\n21:04 patroni detected leader loss\n21:05 db-02 promoted to leader (11s) ✓\n21:05 app error rate 100% — connections still pointing at db-01\n21:07 pgbouncer still routing to db-01, no reconfiguration happened\n21:09 manually restarted pgbouncer, app recovered\n21:09 total outage: 5m 12s (target: 30s)\n21:14 db-01 restarted, rejoined as replica\n21:14 db-01 has 4 transactions db-02 doesn't (async replica was behind)\n21:15 patroni ran pg_rewind, those 4 transactions are gone\n\nthe 4 lost transactions are the part i can't accept. two were payments.", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "pasted-context", "lang": "en"}
{"prompt": "synchronous commit to at least one replica so a promoted leader can't be missing committed transactions", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "synchronous_commit is 'local' on the primary", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "core", "lang": "en"}
{"prompt": "pgbouncer reconfiguration on failover so connections follow the new leader without a manual restart", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "review our connection path end to end and tell me every component that needs to know the leader changed", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "write the failover runbook based on what actually happened, including the pgbouncer step and the data loss check", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "plan the failover testing cadence — monthly, in staging with production-like traffic, with a defined pass criteria", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "database status page for us internally: replication lag per replica, the current leader, and the last failover", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "the lag figure is in bytes and nobody knows whether 4MB is bad", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "boundary", "lang": "en"}
{"prompt": "our app has the database host in 6 places — env vars, a config file, two secrets and a hardcoded default. one source, same connections", "purpose": "refactor", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "after the failover, two services reconnected to the old primary as it came back up as a replica and started getting read-only errors", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "explain what our application does on a read-only error from the database — retry, fail, or crash", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "connection handling that detects a read-only primary and re-resolves the leader rather than erroring for minutes", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "backup verification — restore the latest backup to a scratch instance nightly and assert on a known row and the row counts", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "the restore test passes on an empty backup because it only checks the exit code", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "backup and recovery documentation: what we back up, the rpo, the restore procedure with commands, and the verification", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "plan the point-in-time recovery setup with wal archiving, and how far back we can go", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "wal archiving to s3 with the archive command monitored so a failure is loud", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "archive_command is 'cp' to a local disk that filled up", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "core", "lang": "en"}
{"prompt": "plan the ha architecture then implement the pgbouncer failover handling", "purpose": "planning", "secondary": "backendImpl", "mixed": true, "difficulty": 0.9, "slice": "mixed", "lang": "en"}
{"prompt": "work out how we lost those four transactions and then write the incident report", "purpose": "debugging", "secondary": "writing", "mixed": true, "difficulty": 0.9, "slice": "mixed", "lang": "en"}
{"prompt": "carry on", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "vague-eval", "lang": "en"}
{"prompt": "voice calls. we want to add programmable voice on top of the sms gateway — inbound routing to a sip endpoint or a webhook-driven ivr, recording, and transcription. i want the design including where media flows and what we do and don't want to be in the path of", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "inbound call routing driven by a per-number webhook that returns instructions, with a timeout and a sensible default action", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "call flow builder ui: blocks for play, gather, dial and record, connected, with a test-call action", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "the webhook timeout is 30 seconds so the caller hears silence", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "core", "lang": "en"}
{"prompt": "calls occasionally connect with one-way audio and the logs show nothing wrong", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "review our sip signalling for whether we're leaking internal ip addresses in the sdp", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "voice api documentation: the webhook contract, the instruction verbs, the call status callbacks, and the recording availability timing", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "boundary", "lang": "en"}
{"prompt": "call recording with consent handling per jurisdiction, stored encrypted, with an announced beep where required", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "recordings are stored unencrypted in a public-read bucket", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "core", "lang": "en"}
{"prompt": "our recording consent flag is checked in the api and again in the media server with different logic. one check, and report which numbers change behavior", "purpose": "refactor", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "call detail view: the timeline of events, the recording player, the transcript, and the cost breakdown", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "the recording player loads the whole file before playing", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "core", "lang": "en"}
{"prompt": "plan the call quality monitoring — mos estimation, jitter and loss per leg, and how we tell whether a bad call was us or the carrier", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "per-leg quality metrics collected from the media server and attached to the call record", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "explain how we attribute a bad call to a leg today, and whether we have the data to do it at all", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "voice troubleshooting guide for our support team: the symptoms, what each metric means, and when to escalate to the carrier", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "quality dashboard: mos distribution, calls by carrier, and the worst calls with a drill-in", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "音声通話の片方向音声が時々発生する。まず原因の切り分けをお願いしたい", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "ja"}
{"prompt": "number provisioning — search available numbers by area, purchase, configure, and release, with the carrier apis abstracted", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "number management ui: owned numbers, their configuration, monthly cost, and a search-and-buy flow", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "released numbers stay billable for a month because we don't tell the carrier", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "the number search shows numbers that are no longer available by the time you click buy", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "core", "lang": "en"}
{"prompt": "review our number configuration for whether one customer can point a number they don't own at their webhook", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "number provisioning documentation: what's available where, the regulatory requirements per country, and the porting process", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "boundary", "lang": "en"}
{"prompt": "keep at it", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "vague-eval", "lang": "en"}
{"prompt": "here's the carrier's incident report and our own numbers side by side, and i want to understand the discrepancy before i respond to them:\n\nCARRIER SAYS (14:00–16:00 UTC):\n Messages received from you: 412,004\n Accepted: 411,880\n Rejected (invalid destination): 124\n Delivered: 398,220 (96.7%)\n Failed (handset unreachable): 13,660\n\nWE SAY (same window):\n Messages submitted: 438,910\n Accepted by carrier: 411,880\n No response / timeout: 27,030 <- we retried these\n Delivery receipts received: 366,004\n Still pending: 45,876\n\nso 27k of ours got no response, we retried them, and the carrier's accepted count matches our\naccepted count exactly. and we're missing 32k delivery receipts.", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "pasted-context", "lang": "en"}
{"prompt": "receipt ingestion that can't drop a receipt — persist first, process asynchronously, with a dead letter for unparseable ones", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "our receipt endpoint returns 200 before persisting", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "boundary", "lang": "en"}
{"prompt": "retry-on-timeout that uses a client-generated id the carrier honors, so a retry can't create a duplicate submission", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "review our submission retry logic for whether a timeout followed by a success means we sent twice", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "write the response to the carrier laying out the discrepancy with our numbers, specific and non-accusatory", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "reconciliation report comparing our submission log to the carrier's daily file, with the differences categorized", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "plan how we detect a carrier silently dropping messages within minutes rather than at the monthly reconciliation", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "carrier health monitoring: acceptance rate, receipt latency, and delivery rate per carrier per minute, with alerts on a shift", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "the delivery rate alert has a one hour window so we notice an outage 40 minutes in", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "boundary", "lang": "en"}
{"prompt": "carrier comparison view: cost, delivery rate and receipt latency side by side per destination country", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "automatic carrier failover when the acceptance rate drops, with a manual pin so ops can override", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "the failover flapped between two carriers every 30 seconds for an hour", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "explain how our routing decision is made per message and whether it's reproducible after the fact", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "routing decisions recorded per message so we can explain to a customer why their message went via a particular carrier", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "routing documentation for customers: how we pick a route, what they can control, and what we guarantee about delivery", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "boundary", "lang": "en"}
{"prompt": "our cost calculation exists in the routing decision and again in the billing job with different rate tables. one rate source, and report which messages are priced differently", "purpose": "refactor", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "customers are billed a different amount than the cost we showed at send time", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "plan the routing engine then implement the carrier scoring", "purpose": "planning", "secondary": "backendImpl", "mixed": true, "difficulty": 0.9, "slice": "mixed", "lang": "en"}
{"prompt": "review the receipt handler and fix anything obviously broken", "purpose": "review", "secondary": "quickFix", "mixed": true, "difficulty": 0.7, "slice": "mixed", "lang": "en"}
{"prompt": "next then", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "vague-eval", "lang": "en"}
{"prompt": "the design system's figma-to-code loop. designers change a component in figma, tell us in slack, and someone updates the code a week later, so the two have been out of sync since february. i want the plan for closing this loop, including whether we generate code from figma (i suspect not) and how we detect drift", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "drift detection comparing the figma component properties against our code's prop types, reported weekly", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "drift report ui: components with differences, the figma value against the code value, and a link to both", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "the drift report lists every component because the naming conventions differ between figma and code", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "our figma component names use Title Case With Spaces and our code uses PascalCase, so nothing matches. establish one convention and rename on both sides", "purpose": "refactor", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "boundary", "lang": "en"}
{"prompt": "explain how our token export handles a figma variable that references another variable", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "design system contribution process doc: how a designer proposes a change, who reviews, and how it reaches production", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.5, "slice": "core", "lang": "en"}
{"prompt": "token resolution for aliased figma variables, flattened to concrete values in the export with the alias recorded", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "aliased tokens export as the literal string '{color.brand.500}'", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "core", "lang": "en"}
{"prompt": "plan how we version the design system's figma library alongside the code package, so a designer knows which version their file uses", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "visual regression tests comparing our built components against reference renders exported from figma", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "the visual comparison fails on everything because figma renders text with different antialiasing", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "review our component library for components that exist in code but not in figma, and vice versa", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "component inventory doc listing what exists where, with the gaps called out for both teams", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.5, "slice": "core", "lang": "en"}
{"prompt": "plugin that annotates a figma file with which components are implemented in code and at what version", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.7, "slice": "core", "lang": "en"}
{"prompt": "the annotation overlays sit on top of the designs and can't be hidden", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "boundary", "lang": "en"}
{"prompt": "design the drift detection then implement the token comparison", "purpose": "planning", "secondary": "backendImpl", "mixed": true, "difficulty": 0.7, "slice": "mixed", "lang": "en"}
{"prompt": "explain the token pipeline and then document it for the design team", "purpose": "review", "secondary": "writing", "mixed": true, "difficulty": 0.6, "slice": "mixed", "lang": "en"}
{"prompt": "one more", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "vague-eval", "lang": "en"}
{"prompt": "propose the design for our multi-tenant postgres at the next order of magnitude. we're at 4000 tenants in one database, 2TB, and the biggest tenant is 8% of it. i want to know when we have to shard, what the intermediate steps are (partitioning, moving the big tenants out), and the signals that tell us it's time", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "partition the three biggest tables by tenant hash, online, with the application unaware", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "tenant size dashboard: rows and bytes per tenant per table, with the growth rate and the top 20 by size", "purpose": "frontendImpl", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "the size query does a full table scan per tenant and takes 40 minutes", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "boundary", "lang": "en"}
{"prompt": "queries that were fast are now slow after partitioning, specifically the ones that don't filter by tenant", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "our cross-tenant admin queries don't filter by tenant, which is now a partition scan. rewrite them to be partition-aware or explicitly accept the cost, documented", "purpose": "refactor", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "read our queries and tell me which ones would still work if the tables were on separate databases", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "database scaling documentation: the current architecture, the partitioning scheme, the signals for the next step, and what shards would mean", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "boundary", "lang": "en"}
{"prompt": "tenant move procedure — copy a tenant's rows to another database, verify, cut over with a short write pause, and clean up", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "the tenant move verification compares row counts only, so a mangled column would pass", "purpose": "quickFix", "secondary": null, "mixed": false, "difficulty": 0.4, "slice": "boundary", "lang": "en"}
{"prompt": "a tenant move left rows in both databases and the app read from both depending on the request", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.9, "slice": "core", "lang": "en"}
{"prompt": "explain how our connection routing would decide which database a tenant lives in, if we had more than one", "purpose": "review", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "plan the vacuum and bloat management at this size — autovacuum is falling behind on the biggest tables", "purpose": "planning", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "autovacuum tuning per table with the big ones given more aggressive settings, and monitoring for when it falls behind", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "the events table has 40% dead tuples and autovacuum hasn't run on it in 9 days", "purpose": "debugging", "secondary": null, "mixed": false, "difficulty": 0.8, "slice": "core", "lang": "en"}
{"prompt": "database maintenance runbook: bloat, vacuum, reindexing, and the safe windows for each", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "bloat monitoring with an alert when a table's dead tuple ratio crosses a threshold", "purpose": "backendImpl", "secondary": null, "mixed": false, "difficulty": 0.6, "slice": "core", "lang": "en"}
{"prompt": "design the partitioning then implement it for the events table", "purpose": "planning", "secondary": "backendImpl", "mixed": true, "difficulty": 0.9, "slice": "mixed", "lang": "en"}
{"prompt": "that's the last of it", "purpose": "writing", "secondary": null, "mixed": false, "difficulty": 0.3, "slice": "vague-eval", "lang": "en"}