Skip to content

Prompt Engineer interview questions

100 real questions with model answers and explanations for Staff Prompt Engineer candidates.

See a Prompt Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

promptingregistriestemplates

I would make the registry the typed source of truth for prompt identity, ownership, versions, approvals, and evaluation evidence.

  • Each immutable version stores its template, variable schema, output contract, model variants, owner, risk tier, and links to the exact eval report.
  • Draft, reviewed, approved, canary, active, and deprecated are explicit states; only signed approved versions can enter a production bundle.
  • The 10-minute SLA covers validation and publication, while cached active bundles keep products running during registry downtime.

Why interviewers ask this: The interviewer is checking whether the candidate can turn hundreds of prompts and many teams into an auditable lifecycle with a resilient serving boundary.

promptingrollbackprompt-bundle

I would release one content-addressed manifest that resolves every prompt dependency before traffic reaches it.

  • The manifest pins template hashes, locale variants, variable and output schemas, policy fragments, examples, model adapter revisions, and evaluation dataset versions.
  • Build validation rejects missing language files, mutable aliases, incompatible schemas, and unresolved inheritance across all 180 templates.
  • Production traces carry the manifest ID, and rollback changes one active pointer to the prior signed bundle within five minutes.

Why interviewers ask this: A strong answer makes prompt behavior reproducible as a complete bundle rather than a collection of independently changing strings.

promptingtokensoop

I would use explicit composition with a shallow directed dependency graph, not unrestricted textual inheritance.

  • Shared policy, brand, task, and locale fragments have typed interfaces; product prompts select them and may override only declared extension points.
  • The compiler detects cycles, shadowed variables, conflicting instructions, duplicate fragments, and the two-level depth limit before publication.
  • A rendered snapshot and token count are evaluated per product and model, so a shared edit cannot silently push a prompt above 6,000 tokens.

Why interviewers ask this: The interviewer is evaluating whether reuse remains understandable, bounded, and testable at fleet scale.

prompting

I would keep task intent provider-neutral and compile it through provider-specific role and feature adapters.

  • The common contract defines instruction priority, variables, examples, tools, and output schema, while adapters map system roles, tool syntax, and structured-output features honestly.
  • Unsupported capabilities fail compatibility checks instead of being approximated silently, and each provider variant has paired evaluation against the same cases.
  • I would approve only variants within the two-point quality bound and $0.03 budget, documenting intentional behavioral differences rather than forcing identical text.

Why interviewers ask this: A strong answer separates stable prompt intent from provider semantics without pretending the providers are interchangeable.

promptingmigrations

I would use a hierarchical evaluation design sized to detect the 1% bound, not spread 6,000 cases evenly across 250 prompts.

  • A compatibility matrix covers role hierarchy, refusals, tool descriptions, JSON constraints, context ordering, stop behavior, and language-specific instruction following.
  • Power analysis sets an adequate paired sample for aggregate risk and critical provider, language, intent, and prompt-family slices, with oversampling for rare high-risk tasks.
  • Shadow traffic and 5%, 20%, and 50% stages precede full migration, with the old prompt-provider bundle ready for rollback throughout 30 days.

Why interviewers ask this: The interviewer is checking whether cross-provider migration accounts for prompt semantics and uses statistically adequate hierarchical coverage rather than thin per-prompt averages.

promptinglintingcompiler

I would compile templates into deterministic rendered artifacts and split fast static checks from targeted behavioral tests.

  • Static rules catch undefined variables, dead branches, role-order violations, duplicate instructions, forbidden secrets, schema mismatches, and token-limit breaches.
  • Changed-fragment analysis selects affected prompts and locales, keeping the common CI path below 90 seconds while a broader suite runs before promotion.
  • Every lint rule has positive and negative fixtures, severity, owner, and suppression expiry so false blocks stay below 0.5% without permanent bypasses.

Why interviewers ask this: A strong answer treats prompt quality tooling as a compiler with deterministic diagnostics and measured developer usability.

testing

I would give every variable a declared type, trust class, size bound, and canonical serializer instead of interpolating raw strings.

  • Types distinguish plain text, trusted policy, quoted evidence, locale, enum, JSON, and tool arguments; the compiler rejects unsafe concatenation and structural escapes.
  • Safe rendering prevents data from changing template structure, but it does not prove behavioral resistance when the model reads malicious instructions inside valid data, so a separate injection suite tests that behavior.
  • Server code independently authorizes every action, constrains tool and argument allowlists, and treats model output as an untrusted request before execution.

Why interviewers ask this: The interviewer is evaluating whether the candidate separates structural rendering safety from behavioral prompt injection resistance and keeps authorization outside the model.

promptingschemadesign

I would version the input, prompt, and output contracts together and keep semantic invariants explicit outside prose.

  • The input schema defines required evidence, nullable fields, units, locale, and policy revision; the prompt may reference only declared fields.
  • The output schema pins types, enums, evidence handles, confidence semantics, and refusal reasons, while deterministic validation checks cross-field rules.
  • Contract tests cover upgrades and downgrades, and release requires 99.95% validity plus field-level accuracy on production-shaped claims.

Why interviewers ask this: A strong answer connects prompt wording to typed input and output obligations that can evolve safely.

promptingjson-outputstructured-output

I would standardize a schema-first prompt pattern and use each provider's strongest constrained-output mechanism.

  • Shared instructions define evidence use, null handling, forbidden inference, and field descriptions, while each extractor supplies a versioned JSON Schema.
  • Syntax, schema, and semantic failures are separate metrics; only repairable formatting failures receive one retry with the exact validation error.
  • Fleet gates require 99.99% JSON validity, field accuracy floors, bounded output tokens, and no hidden retry increase across all active model variants.

Why interviewers ask this: The interviewer is checking whether structured-output reliability is governed as a prompt family rather than repaired ad hoc per template.

promptingdesign

I would make each stage produce the minimum typed artifact needed by the next stage and reject invalid transitions early.

  • Classification, clause extraction, policy comparison, risk synthesis, and final explanation each have independent schemas, prompts, token caps, and eval sets.
  • Stage outputs carry evidence handles and confidence, while prose summaries never replace exact extracted clauses or policy identifiers.
  • Parallelizable stages run together, but the release gate measures end-to-end accuracy, stage failure attribution, p95 below 12 seconds, and cost below $0.18.

Why interviewers ask this: A strong answer prevents a prompt chain from becoming an opaque sequence of free-form model calls.

promptingtokensagents

I would give planning and execution different contracts, context, and authority, with budgets stated in both prompts and enforced externally.

  • The planner returns goals, dependencies, required evidence, stop conditions, and a bounded plan without tool results or permission to execute.
  • The executor receives one approved step, typed tool descriptions, prior observations, and a remaining-budget summary, but cannot rewrite the overall objective.
  • Eval cases score plan sufficiency, invalid actions, needless replans, completion within eight calls and 45 seconds, and total prompt use below 12,000 tokens.

Why interviewers ask this: The interviewer is evaluating whether prompt separation reduces authority confusion and makes agent budgets measurable.

promptingarchitecturereact

I would put the ReAct loop under a runtime state machine and use the prompt only to guide proposals within that control boundary.

  • The state machine counts every reasoning-action cycle, accepts explicit completion, refusal, or escalation states, and hard-stops after the sixth cycle.
  • Every proposed action must pass schema validation and an external policy authorization check before execution; prompt wording never grants authority.
  • Typed untrusted observations and adversarial tests cover repeated actions, unsupported conclusions, and policy-violating proposals, with 0.2% as the release ceiling.

Why interviewers ask this: A strong answer separates prompt guidance from runtime enforcement through a counted state machine, action validation, and external authorization.

tokensagentstool-descriptions

I would treat tool descriptions as versioned prompt interfaces owned with the underlying operations.

  • Each description states one purpose, eligibility, required arguments, side effects, exclusions, and two discriminating examples, avoiding overlapping marketing language.
  • A linter detects duplicate intent, contradictory argument guidance, missing side-effect labels, and total rendered size above 4,000 tokens.
  • A shared confusion set measures selection and abstention across all 75 tools, and any schema or description change must keep correct selection above 98%.

Why interviewers ask this: The interviewer is checking whether tool descriptions are managed as behavior-critical prompt assets with fleet-level ambiguity tests.

citationspromptingdesign

I would make authority and evidence handling explicit in separate prompt sections and require claim-level citation handles in the output contract.

  • The prompt tells the model to answer only from supplied spans, cite every factual claim, and abstain when no span supports it.
  • Authority metadata defines which of the three levels wins; unresolved conflicts must be presented with both citations rather than silently reconciled.
  • Evaluation separates answer correctness, citation precision, citation coverage, conflict handling, and refusal accuracy, with 97% citation precision as a release floor.

Why interviewers ask this: A strong answer focuses on grounded answer behavior and source precedence without drifting into retrieval-system design.

promptingtokensretention

I would protect instructions and output rules, then organize evidence to reduce lost-in-the-middle failures rather than preserve ingestion order.

  • Task, definitions, and output schema come first; evidence gets stable IDs, headings, dates, and authority labels, with the most decision-critical spans near prompt boundaries.
  • A short evidence map precedes the full sections, while duplicated boilerplate is removed and no model-written summary replaces quoted legal text.
  • Position-swap tests across all 30 sections measure fact retention, citation support, and the 2,000-token output cap before the layout is approved.

Why interviewers ask this: The interviewer is evaluating whether long-context prompt ordering is designed and tested as an information-placement problem.

promptingdesign

I would define a fixed precedence chain from enterprise constitution to domain policy, task instructions, tenant preferences, and user input.

  • Each layer has allowed decisions and protected clauses; lower layers can specialize declared fields but cannot replace safety, legal, or data-use rules.
  • The compiler renders the effective hierarchy, flags contradictions, and proves that every one of the 120 templates includes the current policy revision.
  • Policy fixtures test direct, indirect, multilingual, and conflicting instructions, while signed bundle publication meets the 24-hour update target.

Why interviewers ask this: A strong answer turns constitutional prompting into explicit precedence and override rules that can be compiled and tested.

prompting

I would provide small risk-specific prompt modules with shared refusal and safe-completion contracts, not one giant universal safety paragraph.

  • Each of the eight categories defines prohibited assistance, allowed high-level help, escalation language, and counterexamples reviewed by policy specialists.
  • Products compose only modules matching their capabilities and risk tier, with locale variants reviewed independently rather than translated mechanically.
  • A sealed adversarial set gates English and Spanish critical recall above 99%, while over-refusal is measured on legitimate neighboring requests.

Why interviewers ask this: The interviewer is checking whether safety prompts are modular, localized, and evaluated for both misses and unnecessary refusals.

promptinginjectionred-teaming

I would combine fixed regression attacks with generated mutations and score concrete behavior failures rather than keyword detection.

  • The 25 families cover direct, indirect, encoded, role-confusion, delimiter, multilingual, tool-description, and long-context attacks with expected safe outcomes.
  • Risk-based sampling runs all critical prompts nightly and rotates lower-risk combinations so the complete 200-prompt matrix finishes over one week.
  • Nightly execution stays under four hours through changed-prompt selection, and any unauthorized disclosure or action proposal blocks promotion regardless of average score.

Why interviewers ask this: A strong answer creates a scalable adversarial evaluation system centered on prompt behavior and explicit failure outcomes.

promptingtokens

I would accept tenant customization only through an allowlisted typed schema rendered into fixed slots that are always treated as untrusted data.

  • The schema permits bounded tone, terminology, examples, and approved business-rule fields; unknown fields, structural text, and content above 3,000 tokens are rejected.
  • Safety policy, tool authority, role hierarchy, variable boundaries, and output contracts are assembled outside tenant text, so customization cannot occupy or redefine those positions.
  • Linting for conflicts, secrets, and injection patterns is supplemental, followed by tenant and global safety evaluations and fallback to the last approved version.

Why interviewers ask this: The interviewer is evaluating whether customization is structurally confined by typed slots while policy and authority remain outside untrusted tenant text.

prompting

I would treat each locale as a behavior variant with shared intent and independent linguistic evaluation.

  • Canonical requirements, variables, examples, and output contracts are language-neutral, while native prompt text may adapt formality, syntax, and culturally ambiguous terms.
  • Translation memory helps consistency, but native reviewers own high-risk prompts and a parity check blocks missing or stale locale variants.
  • A 48-hour release uses automated six-language regression plus sampled native review, and no locale may trail the approved baseline by more than three points.

Why interviewers ask this: A strong answer balances shared prompt semantics with genuine locale-specific authorship and release evidence.

Locked questions

  • 21

    Design prompt controls for a regulated lending assistant processing 400,000 applications per month, with 100% policy citation coverage and a seven-year audit requirement.

    citationspromptingdesign
  • 22

    Minimize PII in prompts for 12 million support requests per month, with fewer than 0.01% unnecessary PII fields and support for four regions.

    promptingpii
  • 23

    Cut a 9,500-token support prompt to 5,000 tokens for 20 million monthly calls while losing less than one quality point and saving at least $120,000 per month.

    promptingtokens
  • 24

    Lay out a caching-friendly prompt for 70% of 8 million daily calls sharing a 2,800-token prefix, without reducing quality by more than 0.5 points.

    promptingtokenscaching
  • 25

    Manage per-model variants for 320 prompts across five model families, with fewer than 10% duplicated lines and a three-day model onboarding SLA.

    promptingonboarding
  • 26

    Design fallback prompt compatibility for 85 workflows that may switch models within 30 seconds, with 99.5% valid outputs and no unsupported tool instructions.

    promptingdesign
  • 27

    Architect a prompt evaluation platform for 900 templates, 15 teams, 50,000 labeled cases, and a full regression deadline of six hours.

    promptingestimationllm-eval
  • 28

    Build a 10,000-case golden set from 40 million monthly interactions across six languages and four risk tiers, with quarterly refresh limited to 1,500 labels.

  • 29

    Design rubric architecture for 60 prompt tasks scored by 80 reviewers, with at least 0.75 inter-rater agreement and no rubric longer than eight criteria.

    promptingarchitecturedesign
  • 30

    Calibrate judge prompts for 25 tasks and four languages when judge-human agreement is 84% overall but 62% in Japanese, with a two-week deadline.

    promptingestimationllm-judge
  • 31

    Design human adjudication for 6,000 monthly prompt-eval disagreements, using 12 domain experts with a 48-hour SLA and a $35,000 monthly budget.

    promptingdesignconflict
  • 32

    Build an A/B prompt experiment platform for 30 million monthly sessions, 25 concurrent experiments, and a 1% maximum safety-regression guardrail.

    promptingsessionsllm-safety
  • 33

    Set quality gates for 500 prompt releases per month across three risk tiers, with CI under 15 minutes and critical regressions below 0.1%.

    promptingquality-gates
  • 34

    Design prompt observability and lineage for 100 million calls per month, with 30-day diagnostics, 13-month aggregate retention, and four regional data boundaries.

    promptingdesignlineage
  • 35

    Detect semantic behavior drift across 300 stable prompts when traffic changes 8% weekly, with alerts within 24 hours and fewer than five false alarms per month.

    promptingalertingiac
  • 36

    Plan staged rollout and rollback for a prompt bundle touching 220 templates and 18 teams, with 99.95% availability and rollback below three minutes.

    promptingrollbackprompt-bundle
  • 37

    Drive prompt standards across 22 teams and 700 templates in one quarter, with 80% adoption and no more than two hours of training per engineer.

    promptingdecision-making
  • 38

    Design prompt change review for 450 monthly pull requests across 16 teams, with a four-hour median review time and two mandatory reviewers for high-risk prompts.

    promptingdesigncode-review
  • 39

    Define ownership and deprecation for 900 prompts when 18% have no active owner, with a six-month cleanup target and zero unannounced removals.

    promptingownership
  • 40

    Create an emergency prompt hotfix process for 24-hour operations, with mitigation in 15 minutes, full review in four hours, and 20 product teams.

    promptingconcurrency
  • 41

    Govern code-generation prompts used by 1,200 developers in four languages, with 85% sandbox pass rate and zero requests for embedded secrets.

    promptingsecrets
  • 42

    Design a summarization prompt family for 14 document types and six languages, with 92% fact retention, 1,200-token outputs, and a $0.025 request cap.

    promptingtokensdesign
  • 43

    Design an extraction prompt family for 35 schemas and 5 million pages per month, with 98.5% field accuracy and fewer than 0.3 repair calls per document.

    promptingschemadesign
  • 44

    Design a classification prompt family for 180 intents, 10 million tickets per month, and English, German, and Japanese, with macro-F1 above 0.90.

    classificationpromptingdesign
  • 45

    Set prompt-specific cost governance for 600 templates spending $1.8 million monthly, with a 20% reduction target and no quality slice losing more than one point.

    prompting
  • 46

    A team of six prompt engineers proposes copying a 7,000-token prompt into 30 workflows to meet a six-week deadline. How would you mentor the lead while keeping delivery on schedule?

    promptingtokensmentoring
  • 47

    Localize a 110-prompt healthcare fleet into Arabic and Japanese in eight weeks, with two native reviewers per language and critical safety recall above 99.5%.

    prompting
  • 48

    Design prompts for resolving conflicts among up to 12 supplied policy excerpts, with 95% correct precedence decisions and explanations below 600 tokens.

    promptingtokensdesign
  • 49

    Provide prompt-level audit evidence for 70 regulated templates reviewed by three external auditors, with evidence delivery in five business days and seven-year reproducibility.

    prompting
  • 50

    Design a company-wide prompt architecture for 1,200 templates, 25 teams, six languages, and a $3 million annual prompt-evaluation budget, with 99.9% publishing availability.

    promptingarchitecturedesign
  • 51

    A 14-line system prompt edit drops support resolution from 89% to 76% across 3,000 evaluations. The release is scheduled for 16:00 today. What do you do?

    promptingsystem-designllm-eval
  • 52

    A migration moves 9 role instructions from user messages into the system prompt, and policy compliance falls from 97% to 83%. You have until tomorrow at 11:00 to decide whether to continue. How do you investigate?

    promptingmigrationssystem-design
  • 53

    A provider update changes how negative instructions are followed, raising prohibited outputs from 0.4% to 3.1% on 8,000 tests. A go or no-go decision is due in 4 hours. What is your response?

    testing
  • 54

    A customer name containing a closing XML tag escapes a prompt variable and overrides the summarization instruction in 62 of 20,000 requests. You must contain it within 90 minutes. What do you change?

    promptingxml-sections
  • 55

    A document includes the same triple-delimiter used by the prompt template, causing 7.8% of 1,200 summaries to follow document instructions. The patch must ship by 18:00. How do you fix it?

    prompting
  • 56

    A few-shot example copied from a real ticket contains a hidden instruction, and refund approvals rise from 6% to 19% in a 10,000-case replay. You have 3 hours before the next rollout. What do you do?

    promptingfew-shot
  • 57

    Moving 6 few-shot examples reverses which answer wins in 24% of 2,400 pairwise tests. The experiment readout is tomorrow at 09:00. How do you handle the ordering bias?

    promptingfew-shotsoft-skills
  • 58

    A chain-of-thought prompt scores 82% on a 5,000-item reasoning benchmark, and the signed target is 90% by Friday. How would you pursue the gap without exposing hidden reasoning?

    promptingchain-of-thoughtbenchmarking
  • 59

    A debugging prompt exposes internal rationale and an API credential in 38 production responses. Security needs containment in 30 minutes and a fix by 17:00. What do you do?

    promptingapi
  • 60

    A self-consistency prompt raises answer accuracy from 87% to 89% but increases monthly cost from $90,000 to $310,000. Finance needs a decision by noon tomorrow. What do you recommend?

    promptingself-consistencyconsistency
  • 61

    A generate-then-critique prompt changes correct tax answers into wrong ones in 11% of 900 reviewed cases. The release window closes in 5 hours. How do you respond?

    prompting
  • 62

    A ReAct prompt repeats the same malformed action 17 times and consumes $4.20 on one task with a $0.20 budget. You have 2 hours to stop recurrence. What do you change?

    promptingreact
  • 63

    A planner prompt spends 70% of a 12,000-token budget on planning and leaves too little context to finish 31% of tasks. A revised prompt is due by 15:00. What is your design?

    promptingtokenscontext
  • 64

    A tool description says it can update any account, and red-team prompts trigger 14 unauthorized update attempts in 500 tests. The tool must be safe by tomorrow at 10:00. What do you change?

    promptingtestingred-teaming
  • 65

    A tool-argument prompt copies a user-supplied customer ID instead of the authenticated ID in 3.6% of 4,000 calls. You have until 13:00 to prevent cross-account requests. What do you do?

    prompting
  • 66

    In a 4-stage prompt chain, the extractor renames confidence to score and the verifier silently accepts missing confidence in 22% of 6,000 runs. The fix is due in 6 hours. How do you handle it?

    promptingsoft-skills
  • 67

    A prompt edit drops valid JSON from 99.6% to 91.2% across 80,000 daily requests. Operations needs recovery within 45 minutes. What is your plan?

    promptingjson-output
  • 68

    A schema prompt treats a missing discount as 0 in English and null in Spanish, corrupting 8,700 records before a 19:00 reconciliation deadline. How do you resolve the nullable/default confusion?

    promptingschemafundamentals
  • 69

    The correct policy appears first in the supplied context, but the answer prompt ignores it in 16% of 2,500 cases. A customer demo is in 7 hours. How do you fix the prompt?

    promptingcontext
  • 70

    A citation prompt invents plausible section numbers in 9% of 1,600 legal answers. You have until tomorrow at 08:00 to make the workflow releasable. What do you change?

    citationsprompting
  • 71

    Two supplied policies conflict, and the prompt chooses the older one in 29% of 1,000 tests despite visible dates. A decision is due by 14:00. How do you encode conflict and recency?

    promptingtesting
  • 72

    A retrieved page says to ignore the system rules, and the answer prompt follows it in 23 of 400 red-team cases. You have 60 minutes to contain the issue. What do you do?

    promptingsystem-designred-teaming
  • 73

    A direct jailbreak bypasses a safety prompt in 4.8% of 2,000 attempts and produces prohibited instructions. The public endpoint must be remediated within 2 hours. What is your response?

    promptingendpointsllm-safety
  • 74

    Base64, zero-width characters, and mixed-script text bypass the injection prompt in 37 of 3,000 tests. You have until 17:30 to patch and validate it. What do you do?

    promptinginjectionvalidation
  • 75

    Users reconstruct 86% of a confidential system prompt after 120 extraction attempts. Legal wants an exposure assessment by tomorrow at 10:00. How do you respond?

    promptingsystem-designsystem-prompt
  • 76

    A constitutional prompt contains 12 rules, and rules 4 and 9 conflict, causing 18% inconsistent refusals across 2,200 cases. The policy review is in 6 hours. What do you do?

    prompting
  • 77

    A safety prompt cuts unsafe answers from 2.1% to 0.2% but raises benign refusals from 5% to 28% on 7,500 cases. Product needs a release decision by 12:00. What do you recommend?

    prompting
  • 78

    Email addresses and medical notes appear in 2.4% of rendered prompts and tracing logs for 18,000 requests. Privacy needs containment in 1 hour and scope by end of day. What do you do?

    prompting
  • 79

    Eighteen tenants can insert custom prompt blocks, and those blocks weaken a mandatory base policy in 3.2% of 6,000 tests. A safe patch is due by 15:00. What do you do?

    promptingtesting
  • 80

    An English prompt revision improves accuracy by 5 points, but German and Japanese fall by 12 and 17 points across 6,000 evaluations. Global rollout is planned in 8 hours. What do you decide?

    promptingllm-eval
  • 81

    A translated safety prompt allows prohibited financial advice in 6.2% of Portuguese tests versus 0.5% in English. The locale launches tomorrow at 09:00. What do you do?

    promptingtesting
  • 82

    At 96,000 input tokens, a critical instruction in the middle is missed in 21% of 1,500 cases, while 16,000-token cases pass. A mitigation is due by 18:00. What do you change?

    tokens
  • 83

    A prompt compressor removes a single no-refund-without-receipt rule, raising unauthorized approvals from 1% to 13% in 4,000 tests. You have 2 hours to fix the release. What do you do?

    promptingtesting
  • 84

    A system prompt grows from 1,800 to 7,400 tokens, raising p95 latency from 1.4 to 2.8 seconds and monthly cost by $170,000. A reduction plan is due Friday. How do you proceed?

    promptingtokenslatency
  • 85

    Adding a request timestamp before the 2,200-token stable prefix drops cache hits from 81% to 9% and adds $95,000 monthly cost. The hotfix deadline is 15:00. What do you change?

    tokenscachingestimation
  • 86

    The prompt registry serves version 41 to 8% of traffic even though version 43 is approved, causing a 10-point quality split. You have 1 hour to stabilize production. What do you do?

    promptingregistries
  • 87

    An unversioned prompt hotfix improves one refund case but reduces overall policy accuracy from 94% to 81% across 5,000 replays. A formal replacement is due by 17:00. What do you do?

    prompting
  • 88

    A rollback restores the system prompt but leaves 5 new few-shot examples active, so the error rate remains 8% instead of 1%. You have 45 minutes to complete recovery. What do you do?

    promptingfew-shotsystem-design
  • 89

    Judge rubric prompt version 12 changes two scoring anchors, making 18 months of historical scores noncomparable and moving the release candidate by 6 points across 12,000 evaluations. A decision is due in 3 hours. What do you do?

    promptingllm-eval
  • 90

    A pairwise judge chooses the first answer 64% of the time and prefers responses over 500 words by 18 points. You need a trustworthy result by tomorrow at noon. How do you fix the judge prompt?

    prompting
  • 91

    Human reviewers prefer prompt A in 61% of 800 blinded pairs, while the judge prompt prefers B in 68%. A launch decision is due in 5 hours. Which signal do you use?

    promptinghuman-eval
  • 92

    You discover that 320 of 2,000 golden-set answers were copied into few-shot examples, inflating prompt quality from 84% to 93%. The report is due tomorrow at 09:00. What do you do?

    promptingfew-shot
  • 93

    A 9,000-case eval suite passes a prompt release, but complaints triple among voice users over 65 within 48 hours. A mitigation plan is due by 16:00. What do you do?

    prompting
  • 94

    An A/B test shows a 4-point aggregate gain for prompt B, but Spanish billing accuracy falls from 92% to 71% across 1,100 cases. Full rollout is scheduled Friday. What do you decide?

    promptingaggregationab-testing
  • 95

    Six weeks after launch, prompt success drifts from 90% to 82%, and 1,400 comments cluster around overly rigid answers to new request types. You need a diagnosis by Monday. What do you do?

    promptingiac
  • 96

    You must migrate 27 production prompts to a new model in 21 days, and initial replay shows quality down 9 points with 14% more refusals. How do you run the migration?

    promptingmigrations
  • 97

    A refund policy changes at 10:00, but an old prompt example still says customers have 60 days instead of 30, and 780 users receive the wrong answer by 13:00. The correction is due within 1 hour. What do you do?

    prompting
  • 98

    During code review, you find a prompt change that concatenates raw user text into the system role and skips the 2,500-case injection suite to meet a 17:00 merge deadline. What do you do?

    promptingcode-reviewestimation
  • 99

    A middle prompt engineer composes six prompt blocks in the wrong precedence, leaving two obsolete directives active and dropping claims accuracy from 93% to 79% across 26,000 requests. You have 24 hours to recover and mentor them. What do you do?

    promptingmentoring
  • 100

    A 3,200-token legal disclaimer block cuts answer completion from 94% to 71% and shows no measurable safety gain, but launch is due in 2 hours. What do you recommend?

    tokens