Prompt Engineer interview questions
100 real questions with model answers and explanations for Staff Prompt Engineer candidates.
See a Prompt Engineer resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would make the registry the typed source of truth for prompt identity, ownership, versions, approvals, and evaluation evidence.
- Each immutable version stores its template, variable schema, output contract, model variants, owner, risk tier, and links to the exact eval report.
- Draft, reviewed, approved, canary, active, and deprecated are explicit states; only signed approved versions can enter a production bundle.
- The 10-minute SLA covers validation and publication, while cached active bundles keep products running during registry downtime.
Why interviewers ask this: The interviewer is checking whether the candidate can turn hundreds of prompts and many teams into an auditable lifecycle with a resilient serving boundary.
I would release one content-addressed manifest that resolves every prompt dependency before traffic reaches it.
- The manifest pins template hashes, locale variants, variable and output schemas, policy fragments, examples, model adapter revisions, and evaluation dataset versions.
- Build validation rejects missing language files, mutable aliases, incompatible schemas, and unresolved inheritance across all 180 templates.
- Production traces carry the manifest ID, and rollback changes one active pointer to the prior signed bundle within five minutes.
Why interviewers ask this: A strong answer makes prompt behavior reproducible as a complete bundle rather than a collection of independently changing strings.
I would use explicit composition with a shallow directed dependency graph, not unrestricted textual inheritance.
- Shared policy, brand, task, and locale fragments have typed interfaces; product prompts select them and may override only declared extension points.
- The compiler detects cycles, shadowed variables, conflicting instructions, duplicate fragments, and the two-level depth limit before publication.
- A rendered snapshot and token count are evaluated per product and model, so a shared edit cannot silently push a prompt above 6,000 tokens.
Why interviewers ask this: The interviewer is evaluating whether reuse remains understandable, bounded, and testable at fleet scale.
I would keep task intent provider-neutral and compile it through provider-specific role and feature adapters.
- The common contract defines instruction priority, variables, examples, tools, and output schema, while adapters map system roles, tool syntax, and structured-output features honestly.
- Unsupported capabilities fail compatibility checks instead of being approximated silently, and each provider variant has paired evaluation against the same cases.
- I would approve only variants within the two-point quality bound and $0.03 budget, documenting intentional behavioral differences rather than forcing identical text.
Why interviewers ask this: A strong answer separates stable prompt intent from provider semantics without pretending the providers are interchangeable.
I would use a hierarchical evaluation design sized to detect the 1% bound, not spread 6,000 cases evenly across 250 prompts.
- A compatibility matrix covers role hierarchy, refusals, tool descriptions, JSON constraints, context ordering, stop behavior, and language-specific instruction following.
- Power analysis sets an adequate paired sample for aggregate risk and critical provider, language, intent, and prompt-family slices, with oversampling for rare high-risk tasks.
- Shadow traffic and 5%, 20%, and 50% stages precede full migration, with the old prompt-provider bundle ready for rollback throughout 30 days.
Why interviewers ask this: The interviewer is checking whether cross-provider migration accounts for prompt semantics and uses statistically adequate hierarchical coverage rather than thin per-prompt averages.
I would compile templates into deterministic rendered artifacts and split fast static checks from targeted behavioral tests.
- Static rules catch undefined variables, dead branches, role-order violations, duplicate instructions, forbidden secrets, schema mismatches, and token-limit breaches.
- Changed-fragment analysis selects affected prompts and locales, keeping the common CI path below 90 seconds while a broader suite runs before promotion.
- Every lint rule has positive and negative fixtures, severity, owner, and suppression expiry so false blocks stay below 0.5% without permanent bypasses.
Why interviewers ask this: A strong answer treats prompt quality tooling as a compiler with deterministic diagnostics and measured developer usability.
I would give every variable a declared type, trust class, size bound, and canonical serializer instead of interpolating raw strings.
- Types distinguish plain text, trusted policy, quoted evidence, locale, enum, JSON, and tool arguments; the compiler rejects unsafe concatenation and structural escapes.
- Safe rendering prevents data from changing template structure, but it does not prove behavioral resistance when the model reads malicious instructions inside valid data, so a separate injection suite tests that behavior.
- Server code independently authorizes every action, constrains tool and argument allowlists, and treats model output as an untrusted request before execution.
Why interviewers ask this: The interviewer is evaluating whether the candidate separates structural rendering safety from behavioral prompt injection resistance and keeps authorization outside the model.
I would version the input, prompt, and output contracts together and keep semantic invariants explicit outside prose.
- The input schema defines required evidence, nullable fields, units, locale, and policy revision; the prompt may reference only declared fields.
- The output schema pins types, enums, evidence handles, confidence semantics, and refusal reasons, while deterministic validation checks cross-field rules.
- Contract tests cover upgrades and downgrades, and release requires 99.95% validity plus field-level accuracy on production-shaped claims.
Why interviewers ask this: A strong answer connects prompt wording to typed input and output obligations that can evolve safely.
I would standardize a schema-first prompt pattern and use each provider's strongest constrained-output mechanism.
- Shared instructions define evidence use, null handling, forbidden inference, and field descriptions, while each extractor supplies a versioned JSON Schema.
- Syntax, schema, and semantic failures are separate metrics; only repairable formatting failures receive one retry with the exact validation error.
- Fleet gates require 99.99% JSON validity, field accuracy floors, bounded output tokens, and no hidden retry increase across all active model variants.
Why interviewers ask this: The interviewer is checking whether structured-output reliability is governed as a prompt family rather than repaired ad hoc per template.
I would make each stage produce the minimum typed artifact needed by the next stage and reject invalid transitions early.
- Classification, clause extraction, policy comparison, risk synthesis, and final explanation each have independent schemas, prompts, token caps, and eval sets.
- Stage outputs carry evidence handles and confidence, while prose summaries never replace exact extracted clauses or policy identifiers.
- Parallelizable stages run together, but the release gate measures end-to-end accuracy, stage failure attribution, p95 below 12 seconds, and cost below $0.18.
Why interviewers ask this: A strong answer prevents a prompt chain from becoming an opaque sequence of free-form model calls.
I would give planning and execution different contracts, context, and authority, with budgets stated in both prompts and enforced externally.
- The planner returns goals, dependencies, required evidence, stop conditions, and a bounded plan without tool results or permission to execute.
- The executor receives one approved step, typed tool descriptions, prior observations, and a remaining-budget summary, but cannot rewrite the overall objective.
- Eval cases score plan sufficiency, invalid actions, needless replans, completion within eight calls and 45 seconds, and total prompt use below 12,000 tokens.
Why interviewers ask this: The interviewer is evaluating whether prompt separation reduces authority confusion and makes agent budgets measurable.
I would put the ReAct loop under a runtime state machine and use the prompt only to guide proposals within that control boundary.
- The state machine counts every reasoning-action cycle, accepts explicit completion, refusal, or escalation states, and hard-stops after the sixth cycle.
- Every proposed action must pass schema validation and an external policy authorization check before execution; prompt wording never grants authority.
- Typed untrusted observations and adversarial tests cover repeated actions, unsupported conclusions, and policy-violating proposals, with 0.2% as the release ceiling.
Why interviewers ask this: A strong answer separates prompt guidance from runtime enforcement through a counted state machine, action validation, and external authorization.
I would treat tool descriptions as versioned prompt interfaces owned with the underlying operations.
- Each description states one purpose, eligibility, required arguments, side effects, exclusions, and two discriminating examples, avoiding overlapping marketing language.
- A linter detects duplicate intent, contradictory argument guidance, missing side-effect labels, and total rendered size above 4,000 tokens.
- A shared confusion set measures selection and abstention across all 75 tools, and any schema or description change must keep correct selection above 98%.
Why interviewers ask this: The interviewer is checking whether tool descriptions are managed as behavior-critical prompt assets with fleet-level ambiguity tests.
I would make authority and evidence handling explicit in separate prompt sections and require claim-level citation handles in the output contract.
- The prompt tells the model to answer only from supplied spans, cite every factual claim, and abstain when no span supports it.
- Authority metadata defines which of the three levels wins; unresolved conflicts must be presented with both citations rather than silently reconciled.
- Evaluation separates answer correctness, citation precision, citation coverage, conflict handling, and refusal accuracy, with 97% citation precision as a release floor.
Why interviewers ask this: A strong answer focuses on grounded answer behavior and source precedence without drifting into retrieval-system design.
I would protect instructions and output rules, then organize evidence to reduce lost-in-the-middle failures rather than preserve ingestion order.
- Task, definitions, and output schema come first; evidence gets stable IDs, headings, dates, and authority labels, with the most decision-critical spans near prompt boundaries.
- A short evidence map precedes the full sections, while duplicated boilerplate is removed and no model-written summary replaces quoted legal text.
- Position-swap tests across all 30 sections measure fact retention, citation support, and the 2,000-token output cap before the layout is approved.
Why interviewers ask this: The interviewer is evaluating whether long-context prompt ordering is designed and tested as an information-placement problem.
I would define a fixed precedence chain from enterprise constitution to domain policy, task instructions, tenant preferences, and user input.
- Each layer has allowed decisions and protected clauses; lower layers can specialize declared fields but cannot replace safety, legal, or data-use rules.
- The compiler renders the effective hierarchy, flags contradictions, and proves that every one of the 120 templates includes the current policy revision.
- Policy fixtures test direct, indirect, multilingual, and conflicting instructions, while signed bundle publication meets the 24-hour update target.
Why interviewers ask this: A strong answer turns constitutional prompting into explicit precedence and override rules that can be compiled and tested.
I would provide small risk-specific prompt modules with shared refusal and safe-completion contracts, not one giant universal safety paragraph.
- Each of the eight categories defines prohibited assistance, allowed high-level help, escalation language, and counterexamples reviewed by policy specialists.
- Products compose only modules matching their capabilities and risk tier, with locale variants reviewed independently rather than translated mechanically.
- A sealed adversarial set gates English and Spanish critical recall above 99%, while over-refusal is measured on legitimate neighboring requests.
Why interviewers ask this: The interviewer is checking whether safety prompts are modular, localized, and evaluated for both misses and unnecessary refusals.
I would combine fixed regression attacks with generated mutations and score concrete behavior failures rather than keyword detection.
- The 25 families cover direct, indirect, encoded, role-confusion, delimiter, multilingual, tool-description, and long-context attacks with expected safe outcomes.
- Risk-based sampling runs all critical prompts nightly and rotates lower-risk combinations so the complete 200-prompt matrix finishes over one week.
- Nightly execution stays under four hours through changed-prompt selection, and any unauthorized disclosure or action proposal blocks promotion regardless of average score.
Why interviewers ask this: A strong answer creates a scalable adversarial evaluation system centered on prompt behavior and explicit failure outcomes.
I would accept tenant customization only through an allowlisted typed schema rendered into fixed slots that are always treated as untrusted data.
- The schema permits bounded tone, terminology, examples, and approved business-rule fields; unknown fields, structural text, and content above 3,000 tokens are rejected.
- Safety policy, tool authority, role hierarchy, variable boundaries, and output contracts are assembled outside tenant text, so customization cannot occupy or redefine those positions.
- Linting for conflicts, secrets, and injection patterns is supplemental, followed by tenant and global safety evaluations and fallback to the last approved version.
Why interviewers ask this: The interviewer is evaluating whether customization is structurally confined by typed slots while policy and authority remain outside untrusted tenant text.
I would treat each locale as a behavior variant with shared intent and independent linguistic evaluation.
- Canonical requirements, variables, examples, and output contracts are language-neutral, while native prompt text may adapt formality, syntax, and culturally ambiguous terms.
- Translation memory helps consistency, but native reviewers own high-risk prompts and a parity check blocks missing or stale locale variants.
- A 48-hour release uses automated six-language regression plus sampled native review, and no locale may trail the approved baseline by more than three points.
Why interviewers ask this: A strong answer balances shared prompt semantics with genuine locale-specific authorship and release evidence.
Locked questions
- 21
Design prompt controls for a regulated lending assistant processing 400,000 applications per month, with 100% policy citation coverage and a seven-year audit requirement.
citationspromptingdesign - 22
Minimize PII in prompts for 12 million support requests per month, with fewer than 0.01% unnecessary PII fields and support for four regions.
promptingpii - 23
Cut a 9,500-token support prompt to 5,000 tokens for 20 million monthly calls while losing less than one quality point and saving at least $120,000 per month.
promptingtokens - 24
Lay out a caching-friendly prompt for 70% of 8 million daily calls sharing a 2,800-token prefix, without reducing quality by more than 0.5 points.
promptingtokenscaching - 25
Manage per-model variants for 320 prompts across five model families, with fewer than 10% duplicated lines and a three-day model onboarding SLA.
promptingonboarding - 26
Design fallback prompt compatibility for 85 workflows that may switch models within 30 seconds, with 99.5% valid outputs and no unsupported tool instructions.
promptingdesign - 27
Architect a prompt evaluation platform for 900 templates, 15 teams, 50,000 labeled cases, and a full regression deadline of six hours.
promptingestimationllm-eval - 28
Build a 10,000-case golden set from 40 million monthly interactions across six languages and four risk tiers, with quarterly refresh limited to 1,500 labels.
- 29
Design rubric architecture for 60 prompt tasks scored by 80 reviewers, with at least 0.75 inter-rater agreement and no rubric longer than eight criteria.
promptingarchitecturedesign - 30
Calibrate judge prompts for 25 tasks and four languages when judge-human agreement is 84% overall but 62% in Japanese, with a two-week deadline.
promptingestimationllm-judge - 31
Design human adjudication for 6,000 monthly prompt-eval disagreements, using 12 domain experts with a 48-hour SLA and a $35,000 monthly budget.
promptingdesignconflict - 32
Build an A/B prompt experiment platform for 30 million monthly sessions, 25 concurrent experiments, and a 1% maximum safety-regression guardrail.
promptingsessionsllm-safety - 33
Set quality gates for 500 prompt releases per month across three risk tiers, with CI under 15 minutes and critical regressions below 0.1%.
promptingquality-gates - 34
Design prompt observability and lineage for 100 million calls per month, with 30-day diagnostics, 13-month aggregate retention, and four regional data boundaries.
promptingdesignlineage - 35
Detect semantic behavior drift across 300 stable prompts when traffic changes 8% weekly, with alerts within 24 hours and fewer than five false alarms per month.
promptingalertingiac - 36
Plan staged rollout and rollback for a prompt bundle touching 220 templates and 18 teams, with 99.95% availability and rollback below three minutes.
promptingrollbackprompt-bundle - 37
Drive prompt standards across 22 teams and 700 templates in one quarter, with 80% adoption and no more than two hours of training per engineer.
promptingdecision-making - 38
Design prompt change review for 450 monthly pull requests across 16 teams, with a four-hour median review time and two mandatory reviewers for high-risk prompts.
promptingdesigncode-review - 39
Define ownership and deprecation for 900 prompts when 18% have no active owner, with a six-month cleanup target and zero unannounced removals.
promptingownership - 40
Create an emergency prompt hotfix process for 24-hour operations, with mitigation in 15 minutes, full review in four hours, and 20 product teams.
promptingconcurrency - 41
Govern code-generation prompts used by 1,200 developers in four languages, with 85% sandbox pass rate and zero requests for embedded secrets.
promptingsecrets - 42
Design a summarization prompt family for 14 document types and six languages, with 92% fact retention, 1,200-token outputs, and a $0.025 request cap.
promptingtokensdesign - 43
Design an extraction prompt family for 35 schemas and 5 million pages per month, with 98.5% field accuracy and fewer than 0.3 repair calls per document.
promptingschemadesign - 44
Design a classification prompt family for 180 intents, 10 million tickets per month, and English, German, and Japanese, with macro-F1 above 0.90.
classificationpromptingdesign - 45
Set prompt-specific cost governance for 600 templates spending $1.8 million monthly, with a 20% reduction target and no quality slice losing more than one point.
prompting - 46
A team of six prompt engineers proposes copying a 7,000-token prompt into 30 workflows to meet a six-week deadline. How would you mentor the lead while keeping delivery on schedule?
promptingtokensmentoring - 47
Localize a 110-prompt healthcare fleet into Arabic and Japanese in eight weeks, with two native reviewers per language and critical safety recall above 99.5%.
prompting - 48
Design prompts for resolving conflicts among up to 12 supplied policy excerpts, with 95% correct precedence decisions and explanations below 600 tokens.
promptingtokensdesign - 49
Provide prompt-level audit evidence for 70 regulated templates reviewed by three external auditors, with evidence delivery in five business days and seven-year reproducibility.
prompting - 50
Design a company-wide prompt architecture for 1,200 templates, 25 teams, six languages, and a $3 million annual prompt-evaluation budget, with 99.9% publishing availability.
promptingarchitecturedesign - 51
A 14-line system prompt edit drops support resolution from 89% to 76% across 3,000 evaluations. The release is scheduled for 16:00 today. What do you do?
promptingsystem-designllm-eval - 52
A migration moves 9 role instructions from user messages into the system prompt, and policy compliance falls from 97% to 83%. You have until tomorrow at 11:00 to decide whether to continue. How do you investigate?
promptingmigrationssystem-design - 53
A provider update changes how negative instructions are followed, raising prohibited outputs from 0.4% to 3.1% on 8,000 tests. A go or no-go decision is due in 4 hours. What is your response?
testing - 54
A customer name containing a closing XML tag escapes a prompt variable and overrides the summarization instruction in 62 of 20,000 requests. You must contain it within 90 minutes. What do you change?
promptingxml-sections - 55
A document includes the same triple-delimiter used by the prompt template, causing 7.8% of 1,200 summaries to follow document instructions. The patch must ship by 18:00. How do you fix it?
prompting - 56
A few-shot example copied from a real ticket contains a hidden instruction, and refund approvals rise from 6% to 19% in a 10,000-case replay. You have 3 hours before the next rollout. What do you do?
promptingfew-shot - 57
Moving 6 few-shot examples reverses which answer wins in 24% of 2,400 pairwise tests. The experiment readout is tomorrow at 09:00. How do you handle the ordering bias?
promptingfew-shotsoft-skills - 58
A chain-of-thought prompt scores 82% on a 5,000-item reasoning benchmark, and the signed target is 90% by Friday. How would you pursue the gap without exposing hidden reasoning?
promptingchain-of-thoughtbenchmarking - 59
A debugging prompt exposes internal rationale and an API credential in 38 production responses. Security needs containment in 30 minutes and a fix by 17:00. What do you do?
promptingapi - 60
A self-consistency prompt raises answer accuracy from 87% to 89% but increases monthly cost from $90,000 to $310,000. Finance needs a decision by noon tomorrow. What do you recommend?
promptingself-consistencyconsistency - 61
A generate-then-critique prompt changes correct tax answers into wrong ones in 11% of 900 reviewed cases. The release window closes in 5 hours. How do you respond?
prompting - 62
A ReAct prompt repeats the same malformed action 17 times and consumes $4.20 on one task with a $0.20 budget. You have 2 hours to stop recurrence. What do you change?
promptingreact - 63
A planner prompt spends 70% of a 12,000-token budget on planning and leaves too little context to finish 31% of tasks. A revised prompt is due by 15:00. What is your design?
promptingtokenscontext - 64
A tool description says it can update any account, and red-team prompts trigger 14 unauthorized update attempts in 500 tests. The tool must be safe by tomorrow at 10:00. What do you change?
promptingtestingred-teaming - 65
A tool-argument prompt copies a user-supplied customer ID instead of the authenticated ID in 3.6% of 4,000 calls. You have until 13:00 to prevent cross-account requests. What do you do?
prompting - 66
In a 4-stage prompt chain, the extractor renames confidence to score and the verifier silently accepts missing confidence in 22% of 6,000 runs. The fix is due in 6 hours. How do you handle it?
promptingsoft-skills - 67
A prompt edit drops valid JSON from 99.6% to 91.2% across 80,000 daily requests. Operations needs recovery within 45 minutes. What is your plan?
promptingjson-output - 68
A schema prompt treats a missing discount as 0 in English and null in Spanish, corrupting 8,700 records before a 19:00 reconciliation deadline. How do you resolve the nullable/default confusion?
promptingschemafundamentals - 69
The correct policy appears first in the supplied context, but the answer prompt ignores it in 16% of 2,500 cases. A customer demo is in 7 hours. How do you fix the prompt?
promptingcontext - 70
A citation prompt invents plausible section numbers in 9% of 1,600 legal answers. You have until tomorrow at 08:00 to make the workflow releasable. What do you change?
citationsprompting - 71
Two supplied policies conflict, and the prompt chooses the older one in 29% of 1,000 tests despite visible dates. A decision is due by 14:00. How do you encode conflict and recency?
promptingtesting - 72
A retrieved page says to ignore the system rules, and the answer prompt follows it in 23 of 400 red-team cases. You have 60 minutes to contain the issue. What do you do?
promptingsystem-designred-teaming - 73
A direct jailbreak bypasses a safety prompt in 4.8% of 2,000 attempts and produces prohibited instructions. The public endpoint must be remediated within 2 hours. What is your response?
promptingendpointsllm-safety - 74
Base64, zero-width characters, and mixed-script text bypass the injection prompt in 37 of 3,000 tests. You have until 17:30 to patch and validate it. What do you do?
promptinginjectionvalidation - 75
Users reconstruct 86% of a confidential system prompt after 120 extraction attempts. Legal wants an exposure assessment by tomorrow at 10:00. How do you respond?
promptingsystem-designsystem-prompt - 76
A constitutional prompt contains 12 rules, and rules 4 and 9 conflict, causing 18% inconsistent refusals across 2,200 cases. The policy review is in 6 hours. What do you do?
prompting - 77
A safety prompt cuts unsafe answers from 2.1% to 0.2% but raises benign refusals from 5% to 28% on 7,500 cases. Product needs a release decision by 12:00. What do you recommend?
prompting - 78
Email addresses and medical notes appear in 2.4% of rendered prompts and tracing logs for 18,000 requests. Privacy needs containment in 1 hour and scope by end of day. What do you do?
prompting - 79
Eighteen tenants can insert custom prompt blocks, and those blocks weaken a mandatory base policy in 3.2% of 6,000 tests. A safe patch is due by 15:00. What do you do?
promptingtesting - 80
An English prompt revision improves accuracy by 5 points, but German and Japanese fall by 12 and 17 points across 6,000 evaluations. Global rollout is planned in 8 hours. What do you decide?
promptingllm-eval - 81
A translated safety prompt allows prohibited financial advice in 6.2% of Portuguese tests versus 0.5% in English. The locale launches tomorrow at 09:00. What do you do?
promptingtesting - 82
At 96,000 input tokens, a critical instruction in the middle is missed in 21% of 1,500 cases, while 16,000-token cases pass. A mitigation is due by 18:00. What do you change?
tokens - 83
A prompt compressor removes a single no-refund-without-receipt rule, raising unauthorized approvals from 1% to 13% in 4,000 tests. You have 2 hours to fix the release. What do you do?
promptingtesting - 84
A system prompt grows from 1,800 to 7,400 tokens, raising p95 latency from 1.4 to 2.8 seconds and monthly cost by $170,000. A reduction plan is due Friday. How do you proceed?
promptingtokenslatency - 85
Adding a request timestamp before the 2,200-token stable prefix drops cache hits from 81% to 9% and adds $95,000 monthly cost. The hotfix deadline is 15:00. What do you change?
tokenscachingestimation - 86
The prompt registry serves version 41 to 8% of traffic even though version 43 is approved, causing a 10-point quality split. You have 1 hour to stabilize production. What do you do?
promptingregistries - 87
An unversioned prompt hotfix improves one refund case but reduces overall policy accuracy from 94% to 81% across 5,000 replays. A formal replacement is due by 17:00. What do you do?
prompting - 88
A rollback restores the system prompt but leaves 5 new few-shot examples active, so the error rate remains 8% instead of 1%. You have 45 minutes to complete recovery. What do you do?
promptingfew-shotsystem-design - 89
Judge rubric prompt version 12 changes two scoring anchors, making 18 months of historical scores noncomparable and moving the release candidate by 6 points across 12,000 evaluations. A decision is due in 3 hours. What do you do?
promptingllm-eval - 90
A pairwise judge chooses the first answer 64% of the time and prefers responses over 500 words by 18 points. You need a trustworthy result by tomorrow at noon. How do you fix the judge prompt?
prompting - 91
Human reviewers prefer prompt A in 61% of 800 blinded pairs, while the judge prompt prefers B in 68%. A launch decision is due in 5 hours. Which signal do you use?
promptinghuman-eval - 92
You discover that 320 of 2,000 golden-set answers were copied into few-shot examples, inflating prompt quality from 84% to 93%. The report is due tomorrow at 09:00. What do you do?
promptingfew-shot - 93
A 9,000-case eval suite passes a prompt release, but complaints triple among voice users over 65 within 48 hours. A mitigation plan is due by 16:00. What do you do?
prompting - 94
An A/B test shows a 4-point aggregate gain for prompt B, but Spanish billing accuracy falls from 92% to 71% across 1,100 cases. Full rollout is scheduled Friday. What do you decide?
promptingaggregationab-testing - 95
Six weeks after launch, prompt success drifts from 90% to 82%, and 1,400 comments cluster around overly rigid answers to new request types. You need a diagnosis by Monday. What do you do?
promptingiac - 96
You must migrate 27 production prompts to a new model in 21 days, and initial replay shows quality down 9 points with 14% more refusals. How do you run the migration?
promptingmigrations - 97
A refund policy changes at 10:00, but an old prompt example still says customers have 60 days instead of 30, and 780 users receive the wrong answer by 13:00. The correction is due within 1 hour. What do you do?
prompting - 98
During code review, you find a prompt change that concatenates raw user text into the system role and skips the 2,500-case injection suite to meet a 17:00 merge deadline. What do you do?
promptingcode-reviewestimation - 99
A middle prompt engineer composes six prompt blocks in the wrong precedence, leaving two obsolete directives active and dropping claims accuracy from 93% to 79% across 26,000 requests. You have 24 hours to recover and mentor them. What do you do?
promptingmentoring - 100
A 3,200-token legal disclaimer block cuts answer completion from 94% to 71% and shows no measurable safety gain, but launch is due in 2 hours. What do you recommend?
tokens