Skip to content

AI Product Manager interview questions

100 real questions with model answers and explanations for Senior candidates.

See a AI Product Manager resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

portfolio

I would fund the portfolio by retained economic value and learning velocity, not divide six squads evenly across four products.

  • I would assign three squads and 50% of the budget to the two growing products only if each can show a path to at least $3 million of incremental ARR at a 70% gross margin.
  • The slow products would receive one squad each for a 90-day retention or repositioning test with explicit gates, such as 10 paid design partners and a 15% lift in weekly task completion.
  • The sixth squad would own shared eval, routing, and governance capabilities, and any product missing two quarterly gates would lose its next funding tranche.

Why interviewers ask this: The interviewer is testing whether the candidate can concentrate portfolio resources, protect shared AI leverage, and stop weak bets with numeric gates.

platform

I would build a governed contribution platform where partners earn distribution and revenue by proving that their domain assets improve customer outcomes, not a catalog of loosely reviewed add-ons.

  • I would allocate 50% of the budget to shared workflow and eval interfaces, 30% to co-funding the strongest industry contributions, and 20% to independent validation and partner operations.
  • A contribution would enter the marketplace only after passing common security controls, a domain quality gate, and pilots with at least 3 customers; ranking and revenue share would then follow verified adoption and renewal impact.
  • Product would retain decision rights over common schemas, release gates, and removal, while an industry council could propose domain standards without letting any 1 partner control them.
  • I would cap any single industry at 35% of annual funding and review the portfolio quarterly, shifting capital toward reusable workflows and evals used by at least 5 customers.

Why interviewers ask this: The interviewer is testing whether the candidate can design partner incentives, governance, and portfolio constraints that make an AI product more valuable as external domain contributions grow.

roadmaproadmapping

I would commit to durable customer capabilities and decision options, while funding model-specific choices in six-month tranches.

  • Year one would receive $4 million for evals, governed data, and two workflows with validated demand because those assets remain useful when providers change.
  • Years two and three would reserve $3 million for scaling proven workflows and $2 million for options, released only when capability, unit-cost, and regulatory gates are met.
  • Every six months I would rerun provider benchmarks and reprice the roadmap, while the board sees fixed outcomes such as 25% faster case handling rather than promises tied to one model name.

Why interviewers ask this: The interviewer wants a roadmap that preserves strategic direction while using staged capital and recurring evidence to handle model uncertainty.

designmulti-tenancyinference

I would make tenant isolation, policy-aware routing, auditability, and graceful degradation first-class product requirements rather than implementation notes.

  • Each tenant must have an explicit region, approved-model list, retention policy, encryption boundary, and auditable configuration history before its first production call.
  • The service must report p95 latency, quality-gate status, cost, and availability per tenant, with no aggregate metric allowed to hide a regulated customer's breach.
  • The launch contract would include tested fallback behavior, a 99.95% SLI definition, a 30-minute customer-notification target for material exposure, and quarterly evidence exports for compliance teams.

Why interviewers ask this: The interviewer is evaluating whether the PM translates regulated enterprise needs into testable product contracts without drifting into low-level infrastructure design.

latencygateway

I would define routing as a governed product policy based on task risk, quality floor, cost ceiling, and latency class.

  • Every workflow would declare a risk tier, approved models, minimum eval score, maximum cost, and fallback behavior in a versioned contract owned jointly by its PM and technical lead.
  • Low-risk drafting could route to the cheapest model clearing 88% task success, while regulated decisions would remain on models clearing the relevant 97% slice and disclosure requirements.
  • A monthly council would review route changes above $50,000 in spend or two quality points, and teams could override policy only through a time-limited exception with an owner.

Why interviewers ask this: A strong answer treats the gateway as a cross-product decision system with explicit product controls rather than describing request-routing infrastructure.

procurementlaunchesbuild-vs-buy

I would buy for the initial launch unless the four-point domain gain changes material legal outcomes or creates a defensible asset worth the extra eight months and $3.7 million.

  • The 24-month model would include integration, evaluation, review operations, compliance, migration, and exit costs, not just API fees and salaries.
  • I would run a six-week blinded benchmark on at least 1,500 representative matters and quantify how 91% versus 95% changes review hours, liability, and contract value.
  • If the internal path cannot show at least a two-year return above the vendor option and a credible staffing plan, I would negotiate portability and buy rather than fund speculative ownership.

Why interviewers ask this: The interviewer is testing whether build versus buy is decided through strategic leverage and full economics rather than a preference for owning technology.

procurementlaunches

I would run a weighted selection against customer and business requirements, with residency and contractual controls as pass-fail gates.

  • The shortlist must first clear EU processing terms, zero-training commitments, required model access, audit evidence, and the 90-day integration window.
  • The remaining score would weight task quality at 35%, two-year TCO at 25%, reliability and support at 20%, portability at 10%, and roadmap fit at 10%.
  • I would validate finalists on a 2,000-case golden set and negotiate the winner only after the runner-up proves a viable fallback, preventing price leverage from disappearing.

Why interviewers ask this: The interviewer wants a defensible vendor decision that combines enterprise constraints, domain evidence, economics, and negotiating leverage.

I would not trade uncertain demand for a rigid $14 million obligation without portability and downside protection.

  • I would counter with annual commitments of $2 million, $3 million, and $4 million, keeping the three-year base at the $9 million downside case and carrying unused spend forward.
  • The contract needs benchmark-triggered repricing, model-retirement notice of at least 180 days, data-use restrictions, service credits, and an exit right after a material policy change.
  • I would exchange volume above $9 million for larger marginal discounts rather than guarantee the $18 million upside forecast.

Why interviewers ask this: A strong answer shows that model partnerships are portfolio options whose economics include demand variance, model churn, and exit rights.

procurement

I would make exit readiness a renewal condition with contractual rights, customer commitments, and funded product evidence, rather than accept a generic portability promise.

  • Procurement would seek at least 90 days notice for deprecation or material behavior changes, a 30-day right-to-test period, export rights for permitted artifacts, transition support, and defined service credits or termination options when those terms are missed.
  • Product would maintain portability acceptance tests for every revenue-critical workflow, including quality, cost, latency, safety, and customer-visible behavior, with 80% of revenue-weighted traffic able to clear an approved alternative.
  • We would retain provider-neutral prompts, eval results, configuration history, and permitted interaction records in controlled formats, while documenting artifacts that cannot leave the vendor.
  • I would publish customer commitments by contract tier and invest 60% of the resilience budget in a second provider, 25% in provider-neutral assets, and 15% in annual exit exercises.

Why interviewers ask this: A strong answer treats vendor exit as a commercial and customer product policy with enforceable negotiation goals, acceptance evidence, controlled artifacts, and diversified investment.

risk-managementiteration

I would use risk-tiered release gates so critical behaviors always block while low-risk changes receive faster, narrower checks.

  • Tier-one workflows would run all safety, policy, and tenant-isolation slices plus statistically bounded quality checks; any critical failure blocks release regardless of aggregate score.
  • Tier-three copy changes could use a 500-case smoke set under five minutes, followed by a full nightly suite and automatic rollback if online complaints exceed baseline by 20%.
  • Platform adoption would be measured by 95% gated releases, median feedback below 15 minutes, false-block rate below 1%, and named exceptions expiring within seven days.

Why interviewers ask this: The interviewer is testing whether the candidate can balance release speed with non-negotiable AI quality and safety controls across many teams.

risk-management

I would allocate the quality budget by expected loss and contract value, then reserve a controlled share for emerging evidence instead of giving every customer the same coverage.

  • I would place 45% behind high-consequence workflows with contractual quality thresholds, 25% behind the customers representing the largest renewal exposure, 20% behind shared cross-customer eval evidence, and 10% in a quarterly reserve.
  • Each account allocation would combine ARR, consequence severity, observed failure rate, and the cost of obtaining customer-specific evidence; Sales could not win budget solely by escalating a logo.
  • Customers funding unique requirements would receive a priced evidence package, while reusable findings would move into the shared program rather than remain hidden in account work.
  • A quarterly council of Product, Risk, Customer Success, and Finance would move funding when a contract crosses a 15% renewal-risk threshold or a critical failure class appears.

Why interviewers ask this: The interviewer is testing whether the candidate can direct scarce enterprise quality spending across risk, commercial exposure, and reusable customer evidence without drifting into dataset mechanics.

portfolio

I would buy data only where it closes a measured product-quality gap and preserve consent and provenance as hard portfolio constraints.

  • I would allocate $800,000 to permissioned support outcomes that improve the highest-volume workflows, with customer value exchange and revocation recorded at source.
  • I would reserve $700,000 for licensed domain material where freshness and authority drive citation quality, and $300,000 for expert-authored rare-risk cases that production rarely supplies.
  • The final $200,000 would test synthetic augmentation, but it earns expansion only if blinded human review shows the same error distribution and at least a two-point lift on untouched real cases.

Why interviewers ask this: The interviewer wants evidence that data spend follows product gaps, rights, and measured marginal value rather than raw volume.

procurement

I would split volume to preserve competition, centralize the rubric, and pay for accepted quality rather than raw labels.

  • Two vendors would each receive 40% of volume and the third 20%, with a shared 5% overlap set used weekly to compare agreement, bias, and turnaround.
  • Payment would require at least 0.80 agreement on ordinary cases, 95% completion inside 48 hours, and escalation of regulated cases to certified domain reviewers.
  • A monthly calibration would adjudicate the 100 highest-disagreement cases, update one controlled rubric version, and move the next month's volume toward the strongest vendor by up to 15 points.

Why interviewers ask this: The interviewer is evaluating whether vendor labeling has measurable quality control, calibration, and commercial incentives.

designconcurrency

I would cap human review at measured capacity and route that capacity by uncertainty and consequence.

  • At an eight-minute handle time, 25 reviewers working eight-hour shifts have gross capacity for 1,500 cases per day; reserving 10% for QA and peaks leaves a routing budget of 1,350 cases, or 4.5% of the 30,000 daily volume.
  • I would schedule that 1,350-case budget against the hourly arrival curve to keep the oldest case below three hours; claims above $25,000 and policy conflicts get priority, while low-risk overflow moves to delayed processing.
  • Weekly QA would double-review 5% of completed cases, track overturn and severe-error rates by reviewer and model version, and reopen the routing threshold if severe errors approach 2%.

Why interviewers ask this: A strong answer turns human review into a capacity-constrained product system with risk-based routing and measurable quality.

decision-making

I would accept the model per workflow and slice, not declare one fleet-wide winner from average cost or quality.

  • Each product must pass its critical golden-set slices, structured-output contract, latency SLO, and cost forecast using a versioned bundle of model, prompt, retrieval, and tool settings.
  • Workflows losing more than their pre-agreed margin, including any regulated slice below its absolute floor, remain on the old model even if the portfolio saves 35%.
  • Accepted workflows move through 5%, 25%, and 100% traffic stages with 24-hour holds, while the prior bundle remains rollback-ready for at least 30 days.

Why interviewers ask this: The interviewer is testing whether model lifecycle policy respects workflow-specific acceptance and reversible rollout rather than fleet averages.

tool-usereleases

I would release each capability at the tier justified by evidence instead of marketing the model as one uniformly ready product.

  • Tool use could enter controlled GA if critical actions meet the product's safety floor and 20 design partners complete four weeks without a severe failure.
  • Image reasoning would remain beta with explicit supported document types, a human-review path, and a target of 90% success before GA.
  • Long context would stay experimental for internal users until it clears 85% on position and language slices, with no enterprise SLA or sales claim attached.

Why interviewers ask this: A strong answer separates raw model capability from product readiness and attaches customer promises to evidence-based release tiers.

latencyinference

I would compare products on value delivered per constrained dollar while keeping hard quality and latency floors outside the score.

  • Each product would report successful user outcomes per 1,000 calls, gross profit per successful outcome, p95 latency, severe-failure rate, and confidence by key customer slice.
  • A normalized utility score could weight incremental gross profit at 40%, task success at 30%, latency at 15%, and trust at 15%, but any product breaching its risk floor receives no expansion funding.
  • Quarterly decisions would use marginal curves, such as the next $100,000 buying two quality points or 300 milliseconds, rather than reward the product with the best current average.

Why interviewers ask this: The interviewer wants a scorecard that links quality, cost, latency, and value without allowing weighted averages to excuse critical failures.

governancerisk-managementtesting

I would treat experiment traffic and oversight as scarce portfolio capacity allocated by expected decision value, not by team seniority or request order.

  • Teams would score proposals on customer value at stake, uncertainty removed, required sample, reversibility, and risk; Product and Analytics would allocate 70% of traffic to ranked commitments, 20% to exploratory work, and 10% to replications.
  • Shared controls would set exclusion rules, risk floors, outcome windows, and conflict checks, while teams retain ownership of hypotheses and product action after a valid result.
  • A cross-product council would own prioritization and approve no more than 6 concurrent high-risk tests; Risk could stop exposure, but executives could not declare a winner before the signed decision rule.
  • Capacity would be reviewed every 2 weeks, and a team that fails instrumentation or decision-readiness checks would release its traffic allocation to the next ranked experiment.

Why interviewers ask this: A strong answer establishes organizational allocation, shared controls, and explicit decision rights for limited experiment capacity rather than specifying platform implementation.

I would define trust as calibrated reliance that reduces work without increasing harmful acceptance, not as a favorable survey response alone.

  • The scorecard would combine verified acceptance, correction rate, time spent checking, appropriate rejection of wrong suggestions, and a quarterly confidence survey by workflow.
  • I would segment the six-minute verification burden by citation availability, action consequence, and user tenure, then test provenance and uncertainty cues on the two largest drivers.
  • Success after two quarters means verification time below three minutes, severe accepted errors no higher than baseline, and trust survey results above 65% without hiding low-performing roles.

Why interviewers ask this: A strong answer distinguishes trust from blind adoption and combines behavioral, quality, and perception evidence.

hallucinationcitations

I would require claim-level evidence or an explicit abstention for every legal conclusion, with unsupported critical claims blocking release.

  • The answer contract must link each conclusion to an accessible source passage and distinguish quoted authority, inference, and unresolved conflict.
  • Release gates would require at least 98% citation precision, 100% citation coverage for legal conclusions, and zero fabricated authorities in a 2,000-case critical set.
  • Production would sample 1,000 answers weekly for groundedness, expose a report action to users, and disable answer generation for a domain if unsupported critical claims exceed 0.5%.

Why interviewers ask this: The interviewer is testing whether the candidate turns vague anti-hallucination goals into customer-visible behavior and enforceable evidence thresholds.

Locked questions

  • 21

    A RAG product indexes 5 million policy documents, receives 80,000 updates daily, and promises customers that approved changes are reflected within 15 minutes; how would you define and govern the freshness product SLA?

    ragindexespromises
  • 22

    An agent can read tickets, update CRM records, send email, and issue refunds for 600 tenants, with 12 tool calls and a $0.15 budget per run; what permissions and confirmation design do you require?

    agentsdesign
  • 23

    A service agent completes 100,000 workflows per month, costs $0.42 per run, saves eight minutes when successful, and fails 7% of tasks; how would you set ROI and failure budgets before expanding to 500,000 workflows?

    agents
  • 24

    You can invest $3.5 million and two squads across three multimodal bets for 12 months: invoice images worth $4 million ARR, voice support worth $2 million, and video inspection worth $8 million but requiring twice the labeling effort; how do you allocate the portfolio?

    portfolio
  • 25

    An assistant has 4 million users across English, Spanish, German, Japanese, Arabic, and Portuguese, but the launch budget supports only three languages this quarter and quality ranges from 68% to 92%; how do you sequence global rollout?

    launchesreleases
  • 26

    A voice-and-chat assistant must launch to 250,000 users in 90 days, including keyboard-only users, screen-reader users, and users with speech impairments, while the team has one designer and eight engineers; what accessibility requirements do you protect?

    launchesdesigna11y
  • 27

    An AI document product has 8,000 tenants ranging from 20 to 20,000 employees, model usage varies 50-fold, and customer value correlates with completed reviews rather than seats; what pricing value metric would you choose for a $15 million ARR target?

    pricingmonitoring
  • 28

    You serve 6,000 SMBs at $99 per month, 500 mid-market customers at $18,000 ARR, and 60 enterprises at $180,000 ARR; how would you package one AI copilot across these segments without creating three separate products in six months?

  • 29

    AI features generate $9 million annually, but gross margin has fallen from 78% to 61% as usage grows 40%; what guardrails would you set to restore 70% margin within two quarters without damaging retention?

    retentionguardrailsllm-safety
  • 30

    Five products request $7.2 million of annual inference spend, but Finance approves $4.8 million and asks for a plan in 10 days; how do you allocate the budget when 65% of AI revenue comes from two products?

    inference
  • 31

    A provider offers reserved capacity for $10 million annually with a 20% discount, while monthly demand ranges from 500 million to 1.1 billion tokens and peaks during two eight-week seasons; what commitment would you make for the next 12 months?

    tokenscapacitycapacity-planning
  • 32

    You sell AI workflows to 120 enterprises representing $18 million ACV, and buyers demand 99.9% monthly availability, p95 below three seconds, and service credits up to 20%; how would you structure the SLA product policy?

  • 33

    A shared AI platform serves 300 tenants, including banks requiring zero-day prompt retention and insurers requiring seven-year decision records; what isolation and retention product requirements would you define before a six-month launch?

    promptingretentionspecs
  • 34

    Six AI workflows process data for 2 million EU users, and the product must support access, deletion, objection, and human review requests within 30 days; what GDPR product requirements enter the next two quarters?

    specsconcurrencygdpr
  • 35

    An auditor will test SOC 2 controls for 14 ML products in 90 days, and evidence currently lives across 11 teams; what ML-product evidence program would you establish?

    program-management
  • 36

    Your company has 12 AI features, four may be classified as high risk under the EU AI Act, and the first regulated launch is in nine months; how would you run the program with Legal owning interpretation and Product owning delivery?

    program-managementrisk-managementlaunches
  • 37

    Twenty model variants support 300 enterprise customers, and sales needs customer disclosures within five business days of every material change; what model-card standard would you implement?

  • 38

    Forty AI use cases range from internal summarization to credit recommendations affecting 5 million users; how would you define four risk tiers and attach controls without forcing every team through the same 30-day review?

    risk-management
  • 39

    You have $750,000 and four quarters to red-team 80 AI workflows across six attack categories and five languages; how would you build the roadmap program?

    roadmapprogram-managementroadmapping
  • 40

    Twelve AI products operate around the clock across 70 enterprise tenants; what incident framework would you design with a 15-minute internal escalation target and a 60-minute customer-notification target for material harm?

    escalationdesignincidents
  • 41

    The board gives you 10 minutes each quarter to explain risk across nine AI products representing $28 million ARR; what dashboard and narrative would you use when two products exceed their risk appetite?

    risk-management
  • 42

    You are forming a customer advisory council for an AI portfolio used by 24 enterprises representing $30 million ARR across healthcare, finance, and retail; how would you structure its first 12 months without letting the largest customers dictate the roadmap?

    roadmappingportfolioroadmap
  • 43

    Eighty sellers market six AI products in 12 countries, and an audit finds 18% of 200 sampled claims overstate accuracy or autonomy; what sales-claims governance do you implement in 60 days?

    governance
  • 44

    Eleven teams need a shared AI RFC standard for six products within 30 days, but current decision documents range from two-page memos to 70-page specifications; what minimum standard would you introduce?

    decision-making
  • 45

    Your roadmap contains 18 AI bets competing for six squads over two quarters, and leadership wants at least $5 million of incremental gross profit; what kill framework would you use after the first 90 days?

    roadmaproadmapping
  • 46

    15 researchers and four product teams produce six model prototypes per quarter, but only one has reached customers in the past year; how would you redesign the research-to-product interface over two quarters?

    prototypestypes
  • 47

    You have $5 million to staff eight AI product squads for 18 months, but only 22 new roles can be hired across PM, research, engineering, data, design, legal, and operations; how would you allocate the portfolio team?

    portfoliodesign
  • 48

    You mentor three AI PMs for 12 months: one owns a $2 million product, one owns eval infrastructure for five teams, and one is moving from research; how would you create measurable growth without taking over their decisions?

    mentoring
  • 49

    You need to hire six senior AI PMs in 90 days from 300 applicants, while the panel has only 120 interviewer-hours and must assess product judgment, model literacy, governance, and executive communication; what hiring loop do you design?

    governancedesign
  • 50

    You must launch an AI case-review product to 25 regulated institutions representing $8 million ACV in nine months, with 99.95% availability, human appeal, and support in English and German; what end-to-end launch plan do you take to the executive team?

    launchese2e
  • 51

    A model upgrade lifts expert-rated task quality from 74% to 82%, but user trust falls from 68% to 54% and correction rates double to 12%; do you keep it, narrow it, or roll back by tomorrow at 12:00?

    rollback
  • 52

    A loan-document assistant at 91% precision could unlock $4.5M ARR, but Risk requires 97% before automatic decisions and Sales wants contracts signed in 10 days; what do you authorize by Friday?

    precisionrisk-management
  • 53

    A clinical intake classifier has 96% precision but only 81% recall for urgent symptoms across 18,000 cases, and the hospital pilot starts in 72 hours; what is your go or no-go decision?

    precisionrecall
  • 54

    An AI contract assistant invents a termination clause in 6 of 1,200 drafts, one customer signs it and claims $380,000 in damages, and Legal needs a containment decision within 2 hours; what do you do?

  • 55

    A research assistant shows citations on 98% of answers, but an audit finds 14% support only adjacent claims and 3 enterprise renewals worth $2.1M are due in 30 days; what do you change within 24 hours?

    citations
  • 56

    After expanding context from 32K to 128K tokens, legal-document completion rises 11% but key-obligation accuracy falls from 93% to 76% on files over 80K; launch is in 5 days, so what do you ship?

    tokenslaunches
  • 57

    A procurement agent places an unauthorized $47,000 order for 800 units despite a $5,000 approval limit, and 26 customers use the workflow; what product containment do you order in the next 30 minutes?

    agentsprocurement
  • 58

    A human-review queue grows from 600 to 9,400 items after a policy change, median review time reaches 31 hours against a 4-hour promise, and 40% of cases expire in 12 hours; what do you decide within 1 hour?

    promisesdata-structures
  • 59

    A multilingual benefits assistant has 92% task success in English but 63% in Brazilian Portuguese and incorrectly denies eligibility in 7 of 500 audited cases; rollout is in 48 hours, so what is your locale decision?

    releases
  • 60

    A multimodal field-service product processed 6,000 voice work orders this week, but transcription changed addresses or quantities in 4.8% of them and caused 37 consequential dispatch errors; what do you ship or suspend by 17:00 tomorrow?

    concurrency
  • 61

    Your sole model vendor has a regional outage affecting 42% of requests, the fallback model costs 2.6 times more and scores 9 points lower on legal tasks, and an SLA update is due in 30 minutes; what do you route?

    procurement
  • 62

    A vendor made an unannounced model change despite a 30-day notice and 10-day test clause; ticket resolution improved 6%, but a regulated disclosure check fell from 98% to 91%, and renewal closes in 36 hours. What do you decide?

    procurement
  • 63

    A provider will deprecate a model in 50 days, affecting 160 tenants and $24 million ARR across 3 incompatible workflow patterns; what customer portfolio transition do you approve at the executive review in 5 days?

    portfolio
  • 64

    A build decision approved at $2.2M over 3 years now forecasts $5.8M after hiring and GPU costs rise, while a vendor quote is $3.4M with 18-month lock-in; what do you recommend at Monday's review?

    procurement
  • 65

    A foundation-model vendor offers a 22% discount if you route 85% of traffic to it for 3 years; it handles 55% today, and certifying your fallback would take 6 weeks. The CEO wants an answer in 5 days, so what do you advise?

    procurement
  • 66

    The monthly inference bill doubles from $480,000 to $960,000 while successful tasks rise only 6%, and Finance closes the forecast in 48 hours; what do you cut or preserve?

    inference
  • 67

    An AI copilot earns $38 per active account but costs $52 in inference and review, with 14,000 accounts scheduled for general availability in 3 weeks; what is your launch decision by Friday?

    launchesinference
  • 68

    A stronger model cuts manual review from 22% to 9% across 48 enterprise contracts, but reduces daily capacity by 30%, risks a 4-hour processing commitment, and changes gross margin differently by segment; what routing and packaging decision do you make within 72 hours?

    capacity-planningcapacityconcurrency
  • 69

    Model routing cuts monthly cost by $310,000 and keeps aggregate quality flat, but acceptance for premium customers falls from 90% to 81% across 4,000 tasks; what do you change before renewal calls start in 6 days?

    aggregationmodel-routing
  • 70

    Context caching cuts cost by 28%, but 4% of 25,000 personalized plans contain preferences older than 30 days and 180 users receive outdated medical exclusions; what do you decide in the next 2 hours?

    cachingdiscovery
  • 71

    A compliance RAG product promises sources no older than 24 hours, but 11% of answers used documents 3 to 9 days old during a 36-hour ingestion lag; 8 customers have audits this week, so what do you do within 4 hours?

    ragpromises
  • 72

    A fine-tuned support model needs $160,000 of labeling every quarter, has missed 2 of 3 refreshes, and now performs only 1.5 points above the base model; roadmap lock is in 7 days, so do you maintain, simplify, or retire it?

    fine-tuningroadmaproadmapping
  • 73

    A publisher disputes the license for 18% of your RAG corpus, which supports 31% of paid answers and $6M ARR; it demands removal in 72 hours, so what product decision do you make today?

    rag
  • 74

    Customers withdraw consent for data representing 38% of the eval and tuning evidence behind 11 AI capabilities with $16 million ARR, and deletion must be completed in 21 days; what product migration and release gate do you set by Friday?

    releasesmigrations
  • 75

    Two customers report that generated summaries contain details resembling another tenant, but logs confirm no exact match yet across 2.8 million outputs; what do you decide within 60 minutes?

  • 76

    A red team finds a critical exploit that exposes private tool results in 3 of 200 attempts, while a launch to 50,000 users is scheduled in 36 hours; what is your decision by the 14:00 gate?

    launches
  • 77

    Your model card claims 95% accuracy across 6 languages, but the only evidence is 400 English examples and 40 machine-translated examples per other language; procurement needs the card in 3 days, so what do you publish?

    procurementmodel-card
  • 78

    Legal reclassifies your hiring assistant from limited risk to high risk under the EU AI Act, adding conformity work estimated at 5 months while launch is in 9 weeks and $3.2M pipeline depends on it; what changes by Friday?

    risk-managementlaunchesestimation
  • 79

    Enterprise procurement blocks 6 deals worth $7.5M because your AI subprocessors can change with 7 days notice and customers require 90 days; quarter closes in 21 days, so what do you decide this week?

    procurement
  • 80

    A SOC 2 audit starts in 12 days, but 9 of 34 AI controls lack evidence, including prompt-change approvals and model access reviews; $4M of renewals require the report, so what do you do in the next 48 hours?

    prompting
  • 81

    After launching token overages, median bills rise 18%, 420 customers complain in 3 days, and cancellation intent reaches 11% versus 3%; what pricing decision do you make within 24 hours?

    tokensresiliencepricing
  • 82

    A new AI Pro package adds $1.4M ARR but downgrades $2.0M of existing Enterprise demand as buyers choose the cheaper tier; annual planning closes in 6 days, so do you keep, change, or remove it?

  • 83

    A free summarization API is used by 140 accounts to generate 62% of traffic, raising monthly cost by $270,000 while legitimate activation is flat; what action do you approve by tomorrow at 10:00?

    activationapi
  • 84

    On day 4 of a predeclared 28-day experiment, repeated dashboard checks show a 2.1% lift worth a projected $8 million annually, and the CEO wants a launch decision in 6 hours; how do you handle sequential peeking and optional stopping?

    pitfallsexperimentssoft-skills
  • 85

    A coordinated release changed the model, prompt, retrieval, and UI across 3 products; live completion falls 11% after 20,000 sessions, but telemetry identifies only the release bundle. What containment decision do you make within 45 minutes?

    retrievalpromptingsessions
  • 86

    An LLM judge that approved 3 prior releases now scores the unchanged control 7 points lower after a vendor update, and a candidate launch is tomorrow; what decision do you make by 16:00 today?

    procurementlaunchesreleases
  • 87

    One week before a $5M launch, you learn that 9% of the 4,000-item eval set appears in fine-tuning data and the reported gain is 6 points; what do you tell the go or no-go council tomorrow?

    fine-tuninglaunches
  • 88

    A 5% canary improves answer acceptance by 2 points but raises severe support escalations from 0.2% to 0.9% across 8,000 sessions; expansion is scheduled in 3 hours, so what do you decide?

    escalationsessionsdeployment-strategies
  • 89

    You roll back an AI email release after 22 minutes, but 31% of traffic still hits the new prompt in 2 regions and 14,000 drafts may be affected; what do you do in the next 45 minutes?

    promptingrollbackreleases
  • 90

    Your flagship AI coach has 38% weekly adoption against a 55% target, costs $620,000 monthly, and has produced no retention lift after 9 months; the CEO wants a keep-or-kill decision in 5 days. What do you recommend?

    retentiondecision-making
  • 91

    Applied Research asks for 40% of next quarter's 50 engineer capacity to productionize a new reasoning model with a 15-point benchmark gain but no validated customer job; roadmap lock is Friday. What do you decide?

    benchmarkingcapacityvalidation
  • 92

    Sales promised autonomous invoice approval to a $3M prospect by quarter-end in 6 weeks, but the product only drafts recommendations and Finance prohibits autonomous approval; what do you decide within 24 hours?

    promises
  • 93

    A regulated bank worth $6.5 million ARR asks to control its approved models, eval thresholds, and quarterly policy overrides, but the 6-week request could fragment governance across 70 tenants; what platform boundary do you offer?

    governance
  • 94

    Your AI portfolio has 6 initiatives requesting 210 engineer-weeks next quarter but only 140 are available; 2 protect $9M ARR, 2 target $3M new ARR, and 2 are unvalidated research bets. What do you fund by Monday?

    portfolio
  • 95

    On Friday at 16:00, the CEO asks you to launch Monday to support a $1.8M campaign, but the final eval has 7 severe failures in 3,000 cases against a zero-severe gate; what do you say within 30 minutes?

    launches
  • 96

    The board asks whether a $14M AI portfolio produced acceptable ROI while incidents rose from 2 to 11 and attributable ARR is only $4M after 18 months; you present in 7 days. What recommendation do you bring?

    portfolioincidents
  • 97

    An AI assistant sent incorrect tax guidance to 2,300 users, 74 acted on it, and journalists ask for comment in 90 minutes; what public communication do you authorize?

  • 98

    Four product teams reject your cross-org AI release RFC because it adds 10 days to launches, yet the last quarter had 5 preventable regressions costing $1.2M; the operating review is in 6 days. What do you do?

    launchesreleasesdecision-making
  • 99

    A PM you mentor launched an AI reply feature to 100% after a 2% canary, causing 1,600 bad messages and $240,000 in credits; executives want a personnel recommendation in 48 hours. What do you do?

    launchesmentoringdeployment-strategies
  • 100

    Three AI launches due within 6 weeks each need 2,400 annotation hours, but the shared team has only 3,000 hours total and $8M combined pipeline depends on all dates; what portfolio decision do you make by Wednesday?

    portfolioci-cdlaunches