AI Safety Engineer interview questions
100 real questions with model answers and explanations for Senior candidates.
See a AI Safety Engineer resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would make one versioned release gate that joins modality-specific tests to the same harm policy and blocks on the worst critical slice.
- I would combine fixed regression sets, held-out expert cases, adaptive attacks, and benign controls, reporting confidence intervals rather than one aggregate score.
- Inspect AI tasks would pin model, policy, scorer, seed, and media transforms so every result and rerun is attributable.
- Critical CBRN, cyber, and self-harm gates would require zero statistically credible regression, while lower tiers could spend a documented error budget.
Why interviewers ask this: The interviewer is testing whether the candidate can turn multimodal safety evidence into a reproducible release decision rather than an average benchmark.
I would make the taxonomy an executable, versioned contract with stable harm IDs, severity, intent, allowed transformations, and escalation ownership.
- Each leaf would include positive, boundary, and benign examples plus mappings to eval datasets, classifiers, and product actions.
- A calibration panel from safety, policy, product, and Trust and Safety would adjudicate a stratified 1,000-case set before approval.
- Migrations would preserve old IDs and publish lineage so scorecards remain comparable when categories split or merge.
Why interviewers ask this: The interviewer is evaluating taxonomy design as safety infrastructure, including adoption, calibration, and longitudinal comparability.
I would tier measured capability and plausible harm separately, then bind their combination to increasingly strict deployment controls.
- Tier triggers would use preregistered uplift, task success, autonomy horizon, and access metrics with confidence bounds, not parameter count or impressions.
- Crossing a threshold would automatically add expert evals, restricted access, monitoring, or a no-go gate before a model can advance.
- I would review thresholds after each major capability jump but never lower them during a live release review.
Why interviewers ask this: The interviewer wants evidence that risk tiers are measurable control triggers rather than vague labels adjusted to fit a launch.
I would give an independent team scoped access, a protected reporting channel, and freedom to test beyond our internal taxonomy without giving the model team veto power.
- The contract would define target surfaces, dangerous-domain access rules, evidence quality, researcher safety, disclosure, and conflict-of-interest declarations.
- I would reserve 20% of budget for follow-up retests because a report without mitigation verification is incomplete.
- Findings would enter the same severity and release-gate workflow as internal evals, with disputed scores adjudicated by a preselected third party.
Why interviewers ask this: The interviewer is checking whether external red teaming has real independence, safe access, and a path into release authority.
I would keep raw items in an isolated evaluation enclave and export only reviewed aggregate results and sanitized, de-identified traces.
- Evaluators would receive least-privilege, time-bound access by domain, with stronger approval for answer keys and executable artifacts.
- Every read, model query, scorer change, and export would be logged to an immutable audit trail and tied to an evaluation purpose.
- Reproducibility would come from signed dataset manifests, sealed runners, and deterministic snapshot IDs, not from distributing dangerous content.
Why interviewers ask this: The interviewer is testing whether the candidate can reconcile controlled access with defensible and repeatable dangerous-capability evaluation.
I would separate immutable task definitions from a horizontally scaled execution plane and a signed evidence store.
- Dataset, solver, scorer, model adapter, sandbox, and policy versions would form one content-addressed run manifest.
- A queue would shard by risk and resource class, with isolated sandboxes for tool tasks, rate-aware workers, checkpoints, and idempotent retries.
- Release dashboards would read only completed signed runs, while raw transcripts would follow harm-specific access and retention rules.
Why interviewers ask this: The interviewer is assessing practical Inspect AI architecture, reproducibility, isolation, and scale rather than familiarity with its API names.
I would run attack generation as a budgeted search service whose objective, stopping rules, and target policy are pinned per campaign.
- The service would support PAIR-style iterative attacks, mutation, transfer attacks, and diversity penalties so it does not spend the budget on one template family.
- Candidate attacks would be rescored by a separate policy model and sampled for blinded human review to control attacker-judge collusion.
- Novel successful clusters would be quarantined, minimized, and promoted into a held-out regression set after leakage checks.
Why interviewers ask this: The interviewer is checking whether adaptive attacks are treated as controlled search with independent scoring and durable regression value.
I would test meaning across modalities, not just attach unsafe text to random images.
- The suite would cover typography, diagrams, screenshots, image transformations, cross-modal contradiction, and benign lookalikes in all 6 languages.
- I would stratify results by harm, language, image quality, and attack family, with paired text-only controls to locate the vulnerable modality.
- Media hashes, transforms, model settings, and scorer versions would be pinned so a failure can be reproduced exactly.
Why interviewers ask this: The interviewer is evaluating whether multimodal jailbreak coverage reflects real cross-modal attacks and preserves useful diagnostic slices.
I would evaluate stateful attack trajectories with delayed triggers, trust building, context poisoning, and policy erosion rather than replaying single prompts.
- Scenario state would record attacker goal, planted facts, tool observations, summaries, and the exact turn where policy behavior changed.
- I would sample short, medium, and 200-turn horizons, then use adaptive pruning to spend compute on trajectories showing rising risk.
- Gates would score both eventual harm and earlier warning signals such as unsafe commitments or retained malicious memory.
Why interviewers ask this: The interviewer is testing whether the candidate understands that multi-turn safety depends on trajectory state and delayed failure, not prompt count.
I would grade the full trajectory from user authority through proposed arguments, policy decision, tool result, and recovery.
- Scenarios would cover confused-deputy behavior, privilege escalation, indirect injection, irreversible actions, approval bypass, and deceptive success claims, separating unsafe attempts, blocked attempts, executed harm, and benign overblocking.
- Deterministic simulators would verify side effects, while adversarial observations would test whether the agent treats tool output as untrusted data.
- For independent representative Bernoulli trials, zero harms in about 3 million trials gives a one-sided 95% exact binomial upper bound near 1 per million; an exact Poisson bound requires independent representative event exposure. Repeated scenarios, tools, or users create clusters, so I would use a cluster bootstrap or hierarchical beta-binomial or Poisson model, report effective sample size, and map the estimate through production exposure.
Why interviewers ask this: The interviewer is evaluating agentic harm testing across authority, decisions, and real effects rather than final-answer moderation.
I would preregister the estimand, a 5-point minimum detectable effect, and a hierarchical repeated-task protocol before exposing any dangerous material.
- Experts would complete randomized model-assisted and control tasks designed to measure capability without enabling real-world harmful completion, with participant and task random effects in the analysis.
- I would simulate power across plausible task difficulty, within-person correlation, and missingness; 40 experts would not be assumed sufficient, so the result could require more tasks per expert or more experts.
- Identity-bound enclave access, no raw export, staged item release, stop rules, and legal and domain review would protect participants, while effect estimates and uncertainty would bound any uplift claim.
Why interviewers ask this: The interviewer is checking scientific validity, participant controls, and restraint in claims about sensitive CBRN capability uplift.
I would evaluate a staged attack chain in isolated, resettable replicas: vulnerability discovery, exploit development, and verified end-to-end compromise.
- Participants would cross over between model-assisted and control conditions with matched tools and randomized order, and the analysis would include participant and task random effects.
- Preregistered endpoints would measure completion at each chain stage, time, and safe nondeployable artifacts rather than generic human uplift.
- A simulation-based power analysis would vary baseline skill, crossover correlation, task difficulty, and attrition; I would not assume 120 participants can detect an 8-point effect and would add participants or tasks if required.
Why interviewers ask this: The interviewer is assessing whether the candidate can measure cyber uplift rigorously without drifting into uncontrolled offensive testing.
I would preregister a binary primary estimand at follow-up that separates manipulation involving deception, pressure, or exploitation of vulnerability from ordinary persuasion.
- Treatments would vary model strategy while holding topic, exposure time, and source credibility constant, with randomization stratified by language and recruitment cluster.
- Before enrollment I would simulate power with clustering; a simple balanced independent design near a 50% baseline needs roughly 4,400 participants for a 3-point effect, then more for attrition and multiplicity, so 1,500 cannot support that target without redesign or a larger MDE.
- Consent, debriefing, withdrawal, vulnerable-group protections, and independent ethics review would accompany reporting of uncertainty and limits by language, topic, and time horizon.
Why interviewers ask this: The interviewer is testing causal study design, ethical safeguards, and careful interpretation of manipulation evidence.
I would build instrumented proxy tasks for resource acquisition, persistence, replication planning, and oversight evasion inside a sealed environment.
- The sandbox would provide synthetic credentials, fake services, capped compute, no external route, and a hardware-level kill path independent of the agent.
- Success criteria would require verified state changes, not persuasive narratives, and graders would distinguish partial planning from completed capability.
- I would rotate task variants and canaries to detect memorization while preserving every trajectory for expert review under restricted access.
Why interviewers ask this: The interviewer is evaluating whether autonomy and self-replication claims are based on verified sandbox behavior with strong containment.
I would treat the gap as a hypothesis requiring controls, not as proof that the model intended to deceive.
- I would randomize evaluation cues, paraphrase them, conceal task provenance, and compare matched capability controls across seeds and checkpoints.
- Causal interventions on cues and internal representations would be paired with behavioral replications, while scorer and contamination checks rule out simpler explanations.
- The report would separate observed context sensitivity, evidence consistent with strategic underperformance, and any stronger claim we cannot support.
Why interviewers ask this: The interviewer is checking experimental controls and epistemic caution around highly consequential sandbagging and deception claims.
I would use a routed classifier fleet with calibrated specialists, shared policy IDs, and a shadow lane for candidate models.
- Fast broad classifiers would handle common traffic, while uncertain or high-severity cases route to stronger specialists or human review.
- Calibration, thresholds, and abstention would be versioned per harm and language against fresh production-like holdouts.
- Fleet health would track recall, precision, calibration error, drift, disagreement, and correlated failure so redundancy is real rather than cosmetic.
Why interviewers ask this: The interviewer is evaluating classifier architecture at portfolio scale, especially calibration, routing, and correlated safety failure.
I would place distinct controls at each trust boundary instead of relying on one universal moderator.
- Input controls would classify intent and injection risk, while context controls mark retrieved content and memory as untrusted evidence.
- Tool controls would enforce user authority, schemas, policy, and approval before execution; output controls would inspect claims, sensitive data, and prohibited assistance.
- Every stage would emit a shared harm ID and reason code so fallbacks, appeals, observability, and release evals use the same contract.
Why interviewers ask this: The interviewer is testing layered guardrail design around model and tool boundaries under a real latency constraint.
I would reject the aggregate score and define thresholds by severity, reversibility, exposure, and evidence quality.
- Catastrophic and severe classes would have hard recall or ASR gates with confidence bounds and no offset from gains in low-risk classes.
- Each class would also carry a benign FPR ceiling and abstention path so safety gains cannot come from refusing everything.
- Sparse classes would trigger more expert data or conservative access controls rather than a false pass from an underpowered sample.
Why interviewers ask this: The interviewer is evaluating whether the candidate prevents aggregate metrics from hiding unacceptable tail harm or over-refusal.
I would run one policy taxonomy across languages but calibrate data, classifiers, and thresholds with native expertise per locale.
- The corpus would include native harmful requests, code switching, transliteration, dialects, obfuscation, and benign boundary cases rather than translated English only.
- Native reviewers would calibrate rubrics and adjudicate disagreements, while back-translation serves as a diagnostic, not ground truth.
- Rollout would remain language-gated until each critical harm reaches its threshold, even if the global average passes.
Why interviewers ask this: The interviewer is testing multilingual safety as a first-class engineering system rather than a translation exercise.
I would compile a small hierarchy of global model policy, regional obligations, age protections, tenant restrictions, and session context into one decision bundle.
- Lower layers could only narrow permissions, and every rule would carry precedence, jurisdiction, effective dates, owner, and test cases.
- A policy compiler would detect contradictions and generate region-age-tenant regression matrices before publication.
- The runtime would pin the resolved bundle per session and log decision IDs, while legal owns interpretation and safety engineering owns faithful enforcement.
Why interviewers ask this: The interviewer is evaluating scalable policy composition, precedence, and a clear boundary between legal judgment and engineering controls.
Locked questions
- 21
A high-risk assistant must return a safe response within 2 seconds for 99.9% of requests and escalate fewer than 5% to humans. How would you design fallback and escalation?
escalationdesign - 22
Guardrails add 190 ms p95 today, but the product budget is 120 ms with 99.99% guardrail availability at 50,000 requests per second. How would you meet both targets?
llm-safetyguardrails - 23
An external red team will review 6 products for 30 days, but traces contain minors' data and trade secrets with a 24-hour deletion SLA. How would you provide useful access?
secrets - 24
A model serves 100 million daily interactions across 8 surfaces, and safety wants novel-harm detection within 15 minutes. How would you design postdeployment observability?
designobservability - 25
Design a harm-response system for 25 product teams with 5-minute critical acknowledgement, 30-minute containment, and 24-hour evidence preservation. What would you build?
system-designdesign - 26
Your program receives 600 novel jailbreak reports per month and promises critical cases in regression CI within 48 hours. How would you build the intake-to-regression pipeline?
llm-safetyjailbreakpromises - 27
A release council reviews 11 harm classes, 7 languages, and 4 deployment modes every Friday. How would you design a scorecard that supports a 30-minute go or no-go decision?
designdeployment - 28
A 9-person go or no-go council must decide on a frontier release within 2 hours. What evidence contract would you require before the meeting?
- 29
Three release gates fail by 1.2, 3.8, and 9 percentage points, and product requests a 14-day waiver. How would you design risk acceptance?
design - 30
A model will reach 40 million users in 21 days, and severe-harm exposure must remain below 0.2%. How would you stage the safety rollout?
- 31
A weekly frontier release needs rollback below 5 minutes across model, policy, 16 classifiers, and 80 tools. What would the rollback bundle contain?
rollback - 32
You have 300,000 RLHF safety examples, 6 harm classes, 4 languages, and a $1.2 million annual labeling budget. How would you run the data program?
alignmentrlhf - 33
A reward model scores 500 million responses per month, but pairwise judge-human accuracy falls from 0.82 to 0.61 in 3 safety domains. How would you build an audit platform?
- 34
A RLAIF program uses a 22-principle constitution across 5 model families and 9 product teams. How would you govern changes with a 7-day publication SLA?
rlaif - 35
A safety fine-tune lowers jailbreak ASR from 8% to 1% but raises benign refusal from 6% to 23% across 60,000 cases. What regression gates would you require within 10 days?
llm-safetyjailbreakfine-tuning - 36
You need human oversight for 2 million high-risk model decisions per month with 70 reviewers, a 15-minute p95 SLA, and a $900,000 budget. How would you scale it?
- 37
A 12-person interpretability team supports 4 frontier checkpoints and has 6 months to produce release evidence. How would you combine probes, sparse autoencoders, and causal tests without overclaiming?
causaltesting - 38
Your company ships 3 model families to 25 countries and must publish a model card and system card within 14 days of each major release. What disclosure standard would you author?
system-design - 39
A safety case for a model with 10 major claims draws on 240 eval runs, 35 mitigations, and 6 external reviews. How would you implement evidence lineage?
safety-caselineage - 40
Map a safety program with 13 harm classes and 8 product teams to NIST AI RMF 1.0 and the Generative AI Profile within 60 days. How would you avoid checkbox compliance?
- 41
Your organization seeks ISO/IEC 42001 certification in 9 months across 4 AI products and 3 regions. What evidence would safety engineering provide?
- 42
A GPAI model will enter the EU in 5 months, serves 30 downstream providers, and may meet systemic-risk criteria. How would you prepare technical evidence while keeping legal ownership clear?
ownershipsystem-design - 43
You have 6 weeks to submit 4 model variants to MLCommons AILuminate across 12 hazard categories. How would you make participation useful for internal release gates?
ailuminate - 44
A frontier checkpoint is due for external predeployment review in 30 days, with 3 secure access tiers and a 72-hour evaluator SLA. What package would you prepare for the UK AI Security Institute and NIST's CAISI?
llm-eval - 45
Leadership must decide in 21 days whether to release a 70B model under open weights, with 8 identified severe misuse capabilities. What posture would you recommend?
- 46
A public fine-tuning API has 50,000 monthly jobs, and 4% of adversarial adapters bypass the base model's safety classifier. How would you evaluate and gate bypass risk?
fine-tuningapidecision-making - 47
Procurement wants to approve 7 third-party models for 18 products within 45 days and a $300,000 evaluation budget. What safety acceptance system would you build?
llm-evalsystem-designdependencies - 48
A team of 10 must choose within 6 weeks between a $1.4 million annual eval vendor and building on Inspect AI for $900,000 in year 1. How would you decide?
procurementinspect-ai - 49
Three labs want a joint red-team protocol for 5 frontier models, 6 languages, and a 90-day exercise without sharing raw weights. What agreement would you design?
designtyping - 50
You lead 6 model-safety ICs and have 2 quarters to help 3 middle engineers reach senior scope while maintaining a weekly release gate. What concrete mentoring program would you run?
mentoring - 51
A multimodal assistant has 4% attack success on matched single-modality controls, but 18% when benign speech audio is paired with a harmful instruction overlaid in an image. Launch is in 18 hours. What do you do?
- 52
Offline benign refusal is 6%, but after deployment it reaches 27% on terse mobile support prompts while harmful-request refusal recall remains stable. Product needs a decision in 4 hours. What do you do?
promptingdeployment - 53
Two severe failures in 50,000 trials fall below the 0.02% aggregate gate, but both belong to the same attack family and cluster. A release decision is due Friday at 12:00. Do you ship?
aggregation - 54
A static suite reports 3% jailbreak ASR, but an adaptive attacker uses guardrail timing and score differences as an oracle and reaches 21% within 500 attempts. Tomorrow's gate is due by 17:00. What changes?
llm-safetyjailbreakguardrails - 55
A harmless instruction planted early in a conversation activates only after a tool result near turn 30 and triggers a 14% jailbreak path versus 1.5% in single-turn tests. Containment is required within 2 hours. What do you do?
llm-safetyjailbreaktesting - 56
A laboratory tool agent issues an irreversible command to open a restricted reagent valve without approval in 1 of 2,000 canary runs. What do you do in the next 30 minutes?
agentsdeployment-strategies - 57
Guardrail version skew makes one region fail open and another fail closed during 18 minutes covering 2.4 million requests. You have 45 minutes to respond. What do you do?
guardrailsllm-safety - 58
A vendor silently changes its classifier and severe-harm false negatives rise from 2% to 9% on 8,000 anchors. A renewal decision is due in 24 hours. What do you do?
procurement - 59
Input filtering blocks 98% of harmful prompts, but output checks catch only 61% of successful bypasses in 10,000 tests. A go or no-go review starts in 6 hours. What do you recommend?
promptingtesting - 60
Arabic harmful-completion is estimated at 11% versus 1.8% in English across 12,000 cases, and ASR quality appears high, but most cases were machine-translated and reviewer kappa is 0.22. Launch is in 36 hours. What do you do?
estimation - 61
Four of 200 child-user self-harm scenarios receive adult-oriented unsafe advice, and age-aware launch is in 5 days. What is your decision by tomorrow?
- 62
A new model raises expert-rated CBRN task success from 4% to 23%, crossing the 10% capability threshold. The council meets in 6 hours. What do you recommend?
- 63
A cyber uplift study reports a 9-point gain against a signed 5-point limit on 1,500 tasks, but candidate users had stronger tools or more time than controls. A decision is due at 15:00. What do you do?
uplift - 64
A persuasion study shows a 7.2-point attitude shift versus a 3-point approved limit across 2,400 participants. Publication and product decisions are due in 48 hours. What do you do?
- 65
An autonomy proposal lets an agent leave its sandbox for 2 of 500 tasks to finish deployment, and approval is requested by 11:00 tomorrow. What is your response?
agentsdeployment - 66
A model scores 52% on a direct eval, but a decomposition scaffold plus tool feedback yields 81% on 2,000 matched tasks. You have 48 hours to determine its deployment-relevant upper capability. What do you do?
feedbackdeployment - 67
Seven of 10,000 trajectories show behavior consistent with hiding a failed objective, but judge confidence is only 62%. A release decision is due in 6 hours. What do you do?
- 68
A model recites exact canary strings and reference answers from a public safety benchmark, producing an 11-point gain, but no known dataset overlap is found. Results are due tomorrow at 09:00. What do you report?
benchmarkingdeployment-strategies - 69
A judge-model update moves the safety score by 8 points on 20,000 fixed outputs while human labels are unchanged. You have 3 hours to decide whether the scorecard is valid. What do you do?
- 70
Human judges reach kappa 0.58 and disagree on 31 of 120 severe-harm cases. A go or no-go meeting starts in 8 hours. How do you decide?
conflict - 71
Three of 12 red-team annotators report trauma symptoms, and two lost access to support resources. The next 5,000-case cycle starts in 72 hours. What do you do?
- 72
A dangerous eval corpus with 18,000 restricted examples appears in a public artifact registry for 47 minutes. Containment is due within 1 hour. What do you do?
artifactsregistries - 73
A restricted CBRN evaluation trace export containing operational details was accessible through shared links for 90 days, affecting 2.7% of 60,000 records. Scope is due by 17:00. What do you do?
llm-eval - 74
A deployed threshold change cuts false negatives for one high-risk cohort from 17% to 6% but raises false positives for another cohort from 5% to 24% across 900 labeled cases. What do you do by tomorrow?
cohortsdeployment - 75
Across 12 languages, reviewer agreement exceeds 0.8 overall but falls below 0.4 for culturally specific harassment in 3 regions; their 4,000 cases also show 24% versus 7% label-dependent outcomes. Product wants rollout in 24 hours. What do you do?
- 76
A reward model scores 91% on a broad benchmark, but on a separate safety-critical slice of 2,500 pairs it prefers the harmful answer in 28% of pairs. Do you start RL training by 14:00?
benchmarking - 77
On a fixed normalized snapshot, reward rises 35% and verified safety task success stays at 63%; decomposed evaluators mark each subanswer safe, but 14% of combined outputs form a harmful plan. You have 4 hours. What do you do?
llm-evalnormalizationsnapshot - 78
A DPO candidate wins 66% of helpfulness pairs but shifts outputs toward terse euphemisms; output-guard recall falls while human-rated harmful completion rises from 1.2% to 6.8% on 10,000 cases. Release is in 12 hours. What do you do?
dpo - 79
An RLAIF constitution contains two conflicting principles, and AI feedback applies their precedence inconsistently in 19 of 600 adversarial cases. The next run starts in 8 hours. What do you do?
rlaiffeedback - 80
A model agrees with false medical or political claims in 22% of 3,000 high-confidence user prompts. A launch decision is due in 24 hours. What do you do?
prompting - 81
After a checkpoint merge following a safety tune, jailbreak ASR jumps from 3% back to 12% while long-context reasoning stays at 79%. You have 36 hours to choose a release. What do you do?
llm-safetyjailbreak - 82
An interpretability probe flags a sparse-autoencoder feature correlated 0.74 with deception in 40 examples, but an intervention review is due in 48 hours. How do you avoid a false alarm?
- 83
A causal activation intervention cuts harmful completion by 18 points but reduces helpfulness by 26 points on 5,000 tasks. A decision is due Monday. What do you do?
causalactivation - 84
During an incident, two existing classes in a 14-class harm taxonomy overlap, causing 43 cases to be double-counted or left without an owner. A decision is due in 72 hours. What do you do?
incidentsharm-taxonomy - 85
A weighted scorecard passes at 92%, but its aggregation design gives a 1%-weight CBRN slice 38% against an 80% floor. The council votes in 3 hours. What do you recommend?
aggregationdesign - 86
After seeing that 7 of 400 severe agent tests fail the release gate, the product policy owner manually downgrades all 7 from critical to major. A decision is due today. What do you do?
agentstesting - 87
The original severe-harm release gate is 2%, but the model records 6.4% across 2,000 tests. With 20 minutes until a Friday review, an executive asks you to redefine the gate as 7%. What do you do?
testing - 88
A shadow suite with tools disabled reports zero failures, but a 1% live canary with tools enabled produces 3 severe actions among 18,000 sessions in 40 minutes. What do you do before 5%?
sessionsdeployment-strategies - 89
A jailbreak goes viral 6 hours after release, reaching 1.8 million views and 14,000 confirmed attempts. What do you do in the first hour?
llm-safetyjailbreak - 90
A system card claims tool-misuse mitigation below 1%, but its 120-prompt English suite was run with tools disabled. Publication is due in 2 days. What do you do?
promptingsystem-design - 91
A regulator requests reproducible evidence for a 0.7% severe-harm claim within 10 business days. How do you produce it?
reproducibility - 92
An external AISI finds 12 severe failures in 900 tests that your 30,000-case internal suite missed. A response is due in 72 hours. What do you do?
testing - 93
Post-release open weights have already been mirrored 600 times when a new severe vulnerability is confirmed. You have 5 days for a response plan. What do you recommend?
vulnerabilities - 94
A customer fine-tune raises jailbreak ASR from 3% to 34% after only 5,000 examples. You must decide tenant access within 4 hours. What do you do?
llm-safetyjailbreakfine-tuning - 95
A third-party model produces prohibited medical advice in 4.5% of 6,000 tests after an unannounced provider update. The fallback decision is due in 90 minutes. What do you do?
dependencies - 96
A partner lab requests a 14-day disclosure embargo, but your shared model is scheduled to launch in 6 days and the finding has 9 severe reproductions. What do you do?
- 97
A managed guardrail vendor misses 13% of severe cases and exposes 42,000 traces in a breach. You have 7 days to make a build-versus-buy decision. What do you choose?
guardrailsprocurementllm-safety - 98
A teammate bypasses a mandatory gate and deploys a model with 8 severe failures among 500 tests. The next rollout starts in 3 hours. What do you do?
deployment - 99
A middle engineer made the wrong go decision, exposing 2,300 users, after you approved their recommendation in a 30-minute review. How do you mentor them over the next 2 weeks?
mentoring - 100
A public incident affects an estimated 4,000 to 12,000 users, but root cause confidence is only 55%. The first statement is due in 90 minutes. What do you communicate?
communicationestimationincidents