Cybersecurity Analyst interview questions
100 real questions with model answers and explanations for Senior candidates.
See a Cybersecurity Analyst resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would tune the highest-cost detections against labeled outcomes while holding validated ATT&CK coverage constant.
- In Splunk ES, I would rank the 120 searches by analyst minutes and false-positive rate, then tune the top 15 that consume 70% of effort before touching low-volume rules.
- I would add asset criticality, identity context, and 14-day baselines to SPL thresholds, accepting up to 10 minutes of correlation latency to reduce benign alerts by at least 60%.
- I would replay Atomic Red Team tests for the 20 priority ATT&CK techniques after each change and roll back any rule whose recall drops below 95%, even if the 5% false-positive target is missed.
Why interviewers ask this: The interviewer is evaluating whether the candidate can tune a large Splunk estate with measured precision, protected recall, and realistic query costs.
I would correlate low-and-slow failures across identities and sources instead of alerting on a fixed per-user threshold.
- In Sentinel KQL, I would summarize failures over 10 minutes by source IP, tenant, and authentication protocol, requiring at least 20 users and a failure ratio above 95% to suppress ordinary mistypes.
- I would enrich with Entra ID named locations and service-account lists, accepting a reviewed allowlist of no more than 30 scanners to raise precision from the current baseline to 90%.
- I would validate T1110.003 with 12 positive and 40 benign fixtures in the Sentinel test workspace, keeping scheduled-query frequency at 5 minutes so end-to-end latency stays below 15 minutes despite higher compute cost.
Why interviewers ask this: A strong answer combines correct password-spray aggregation, identity context, explicit test data, and a latency-versus-cost decision.
I would detect suspicious shell sequences and context rather than every interactive shell launch.
- In Elastic EQL, I would correlate a network-facing parent process with shell execution and a follow-on utility within 60 seconds, reducing the search from 180,000 EPS to a bounded three-event sequence.
- I would require rare parent-child pairs or execution by service accounts, using a 30-day baseline and accepting a 24-hour learning delay to keep daily volume below 300 alerts.
- I would test 15 T1059.004 variants plus 100 deployment-script samples in a representative 180,000-EPS replay, rejecting the rule if p95 query latency exceeds 2 minutes or malicious-sequence recall falls below 90%.
Why interviewers ask this: The interviewer checks whether the candidate can express a behavioral EQL analytic that remains selective and performant at high event volume.
I would replace label coverage with evidence that telemetry, analytics, and alert delivery work for each prioritized behavior.
- In an ATT&CK Navigator layer, I would record all 40 techniques with data source, platform, rule ID, owner, and last test, counting visibility, detection, and prevention as 3 separate states.
- I would run Atomic Red Team or CALDERA tests for at least 10 techniques per sprint, accepting 2 hours of controlled test-host downtime to complete all 40 within 90 days.
- I would report validated detection rate and test age, requiring 90% of the 40 techniques to alert within their SLO while exposing any technique older than 90 days as unvalidated rather than green.
Why interviewers ask this: The interviewer is testing whether ATT&CK coverage represents current end-to-end evidence instead of self-reported rule mappings.
I would keep Sigma as the portable detection intent while treating each SIEM translation as a tested platform artifact.
- I would store all 180 Sigma rules with logsource, ATT&CK mapping, owner, and 6 positive and negative fixtures, then use pySigma backends to generate SPL, KQL, and Elasticsearch queries.
- I would permit platform overrides for at most 15% of rules where joins or field semantics differ, accepting duplicate maintenance only when native logic improves precision by at least 10 percentage points.
- CI would replay every generated query against 7 days of sampled data and block release below 85% matched outcomes, while a 30-day exception expiry prevents silent long-term divergence.
Why interviewers ask this: A strong answer treats Sigma as portable intent, not guaranteed equivalence, and controls necessary platform divergence with tests and expiry.
I would make versioned tests, scoped shadow evaluation, and Git rollback mandatory for every detection change.
- GitHub Actions would lint 25 weekly rules, validate schemas, and run at least 5 positive and 20 benign fixtures per rule, adding about 8 minutes to CI.
- Splunk REST, Sentinel Content, or Elastic Detection Engine APIs would publish candidates as shadow rules for 7 days only in named test workspaces or indexes, blocking promotion above 125% of baseline volume or below 80% precision.
- Production promotion would deploy the approved Git revision through the same API, and rollback would reapply the previous Git tag within 15 minutes to keep failed changes below 2%.
Why interviewers ask this: The interviewer evaluates whether detection content receives production-grade tests and rollout controls without making delivery impractically slow.
I would retain the valuable behavior but move benign distinctions into explicit, measurable conditions.
- In CrowdStrike Falcon LogScale, I would label 30 days of outcomes by command line, signer, parent, user, and asset tier, then identify the 5 benign clusters responsible for at least 80% of volume.
- I would suppress only signed, approved combinations through an owner-bound allowlist expiring in 45 days, accepting up to 20 minutes of weekly maintenance to reach fewer than 500 alerts.
- I would replay the 2 material cases plus at least 20 adversary-emulation variants after each change and reject tuning if recall falls below 95%, even when analyst workload remains high.
Why interviewers ask this: The interviewer is looking for disciplined false-positive reduction that protects rare high-value signal through labeled data and regression tests.
I would fund fields and events required by the 30 detections before paying to ingest every available record.
- In a telemetry matrix, I would map each of the 30 detections to required fields and rank sources by unique coverage, keeping Entra audit, EDR process, DNS, and cloud control-plane events that support at least 2 priority behaviors.
- I would use Cribl Stream to drop verbose success events and route raw copies to object storage, accepting 30-minute cold retrieval to reduce SIEM ingestion from 9 TB to at most 5 TB per day.
- I would monitor event arrival and required-field completeness with a 98% daily SLO, restoring any filtered class within 4 hours if a canary detection or schema check loses coverage.
Why interviewers ask this: A strong answer ties ingestion cost decisions to named detection requirements and measurable telemetry health rather than arbitrary log reduction.
I would first prove whether the stream fits the link, then reserve enough headroom to drain outage backlog.
- A 200 Mbps link carries 25 MB/s, about 208 bytes per event at 120,000 EPS before overhead; if measured compressed wire size exceeds that budget, full real-time delivery is mathematically impossible.
- Cribl would keep the detection-critical stream below 50% of link capacity, roughly a 104-byte-per-event equivalent at full EPS, and retain raw low-priority events regionally for later transfer or a secondary path.
- Kafka would retain at least 2 hours, with disk sized from measured stored bytes and replication factor; 50% spare bandwidth must drain a 30-minute critical backlog within the next 30 minutes in the outage test.
Why interviewers ask this: The interviewer checks whether the candidate understands buffering capacity, traffic prioritization, and observable failure testing in a security pipeline.
I would correlate identity risk and session properties instead of treating geographic distance alone as credential misuse.
- In Microsoft Sentinel, I would join Entra sign-ins across the 60 apps with device ID, ASN, MFA result, token ID, and 30-day user baselines, requiring 2 inconsistent session attributes beyond location.
- I would model 12 approved VPN egress ranges and mobile carrier ASNs as context rather than blanket exclusions, accepting a 5-minute enrichment delay to target 85% precision.
- I would validate T1078 with 25 replayed anomalous sessions and 200 benign travel samples, keeping recall above 90% while reducing daily alerts from 1,600 to below 250.
Why interviewers ask this: A strong answer uses identity and session evidence to overcome the known weakness of location-only impossible-travel analytics.
I would baseline each service account's allowed hosts, protocols, and schedule, then alert on bounded deviations.
- Splunk UBA would learn 30 days of source hosts and hourly patterns for all 6,500 accounts before enforcing deviations.
- Splunk ES would join 4624 logon types 2 or 10 to 4672 by Logon ID and require a new host or schedule; 4672 means special privileges were assigned, not privileged-group access.
- A separate signal would track membership changes in 4728, 4732, and 4756 and correlate the changed account with a later anomalous logon.
- I would onboard the 200 highest-privilege accounts first and replay 20 approved job changes, targeting under 100 alerts daily and below 5% failed batch jobs.
Why interviewers ask this: The interviewer evaluates whether the candidate can detect service-account misuse while respecting workload regularity and operational change.
I would score domain novelty and tunneling-like query behavior together so common DNS noise does not dominate.
- In Elastic Security, I would aggregate 15-minute Zeek DNS windows by host and domain, scoring label length above 40 characters, entropy, NXDOMAIN ratio, and more than 500 queries as T1071.004 features.
- I would enrich with a 30-day domain baseline and enterprise resolver allowlist, accepting a 24-hour delay for newly approved services to reduce 90 million queries to fewer than 50 cases per day.
- I would test 10 controlled iodine or dnscat2 patterns and 100 software-update domains, requiring at least 90% malicious recall and below 2% false positives in the labeled set.
Why interviewers ask this: A strong answer combines concrete DNS-tunneling features, scale-aware aggregation, and benign software validation.
I would correlate encoded execution with provenance, content, and follow-on behavior rather than suppressing PowerShell broadly.
- CrowdStrike Falcon would join PowerShell command lines with parent process, signer, user, network connection, and AMSI result, requiring at least 2 suspicious features within 5 minutes.
- I would allowlist 40 signed administration scripts by SHA-256 and owner for 30 days, accepting weekly recertification work to reduce 4,500 alerts by at least 80%.
- Atomic Red Team T1059.001 tests would cover 15 encoding and download variants, with release blocked if recall falls below 95% or p95 alert latency exceeds 3 minutes.
Why interviewers ask this: The interviewer checks whether the candidate can tune a common PowerShell analytic without creating a broad signed-script blind spot.
I would detect new AWS role trust paths and keep policy attachment in a separate privilege-expansion analytic.
- Splunk ES would compare old and new trust documents in CloudTrail UpdateAssumeRolePolicy, correlating actor, target role, new trusted principal, source account, and a 45-day baseline.
- It would alert on an unapproved new principal and raise confidence when that principal successfully calls AssumeRole, exempting automation only with an approved identity and change ID to stay under 30 alerts daily.
- AttachRolePolicy and PutRolePolicy would feed a separate privilege-expansion analytic, while 12 trust-change tests and 100 approved deployments must deliver 95% recall within 10 minutes.
Why interviewers ask this: A strong answer distinguishes legitimate infrastructure automation from anomalous privilege changes using specific cloud audit semantics.
I would retire or restructure rules using test evidence and threat value, not alert frequency alone.
- Splunk Monitoring Console would rank all 260 rules by search cost, skipped executions, ATT&CK priority, and validated findings, targeting the 20 searches responsible for 60% of scheduler load.
- I would convert broad real-time searches to 5-minute accelerated data-model queries where the SLO allows it, accepting up to 5 minutes more latency to cut total compute by 25%.
- The 90 quiet rules would remain only if Atomic Red Team tests pass and their technique is prioritized; otherwise an enabled 30-day non-paging shadow search with owner approval would precede retirement.
Why interviewers ask this: The interviewer is evaluating whether the candidate can control SIEM cost while preserving evidence for rare but important detections.
I would calculate capacity before assigning tiers and reduce human-reviewed volume to what the roster can actually handle.
- All 40 people provide at most 8 gross minutes per alert at 12,000 alerts and a 40-hour week; with 30% shrinkage that falls to 5.6 minutes before reserving investigation or engineering capacity.
- If 10 staff remain on complex investigations and 6 on detection engineering, the 24 triage analysts have 4.8 gross minutes, so a new 3-minute handoff is impossible without reducing intake.
- Enrichment and validated risk splitting would cap human review near 5,040 alerts at 8-minute AHT and 30% shrinkage, while ServiceNow evidence fields and weekly QA target rework below 8%.
Why interviewers ask this: A strong answer turns SOC tiers into measurable decision and evidence boundaries while accounting for staffing and skill flow.
I would use MDR after-hours triage with a hard internal on-call budget and contracted overflow.
- The MDR would acknowledge all 180 monthly high-severity alerts within 10 minutes, collect required evidence, and resolve or hand off cases under approved playbooks.
- A six-person senior rotation would accept no more than 6 overnight escalations per week in total; the seventh and later cases would remain with the MDR overflow team until a staffed handoff.
- Each overnight page would trigger 8 hours of recovery and a named replacement for the analyst's next shift, with that lost capacity included in the roster plan.
- A 90-day pilot would measure missed cases, acknowledgment, overflow, recovery compliance, and cost before any move in-house.
Why interviewers ask this: The interviewer checks whether the candidate can satisfy service coverage without hiding unsustainable on-call demand.
I would automate evidence collection and make the analyst decide only the ambiguous classification and scope.
- Cortex XSOAR would retrieve headers, URL detonation, sender history, recipient count, and mailbox actions in under 3 minutes, saving about 15 analyst minutes per each of 700 cases.
- The playbook would auto-close only when 5 benign checks agree and confidence exceeds 98%, accepting manual review for roughly 20% of cases to keep mistaken closure below 2%.
- A 100-case weekly QA sample in ServiceNow would measure median handling and closure error, with any 2-week breach of 12 minutes or 2% reverting the last automation change.
Why interviewers ask this: A strong answer automates repeatable collection while retaining measurable human judgment at the risky decision point.
I would move supported first-pass analytics closer to regional identity telemetry and isolate its transport from bulk logs.
- Regional Sentinel workspaces would use NRT rules on their one-minute cycle for compatible identity logic; unsupported joins would use the platform's supported scheduled minimum, such as 5 minutes, not an imaginary 2-minute schedule.
- Identity events would use a separate Event Hub and dedicated consumer group from bulk telemetry, with 30 minutes of regional buffering; Event Hubs would not be described as prioritizing records inside one stream.
- I would budget the 8 minutes by rule type, such as 2 minutes ingestion, 5 minutes scheduled analysis, and 1 minute delivery, and page the owner on a segment breach.
Why interviewers ask this: The interviewer evaluates whether the candidate can decompose MTTD and change architecture around a fixed network constraint.
I would remove context wait states and reserve analyst effort for decisions that cannot be automated safely.
- ServiceNow SecOps would enrich every case from CMDB, Entra ID, and CrowdStrike within 2 minutes, accepting a daily 1-hour reconciliation job to eliminate most of the current 35% wait.
- Cortex XSOAR would run 6 read-only queries in parallel and attach results to a standard timeline, targeting 30 minutes saved per high-severity case without automated containment.
- I would report p50 and p90 MTTR by playbook for 8 weeks, promoting the change only if median falls below 4 hours and reopened cases stay under 3%.
Why interviewers ask this: A strong answer improves MTTR through concrete workflow latency removal rather than pressuring analysts to close cases faster.
Locked questions
- 21
An on-call team of 8 receives 320 pages per month, but only 25% require action; how would you bring the 30-day average below 1 actionable interruption per staffed analyst-shift without weakening P1 coverage?
on-callcoverage - 22
An MSSP monitors 14 subsidiaries but sends 30% of escalations to the wrong local team; how would you design routing and evidence handoffs to get misroutes below 5% within 60 days?
escalationdesignmonitoring - 23
Quarterly exercises show that a 15-person SOC can handle 2 simultaneous P1 cases but fails at 5; how would you build surge capacity for a 6-hour peak without hiring this quarter?
soc-operationscapacity - 24
A vulnerability program has 80,000 assets and 1.2 million findings, but only 10 engineers can coordinate remediation; how would you use CVSS, EPSS, CISA KEV, and asset value to select the first 5,000 findings?
vulnerabilitiesvuln-management - 25
You must define patch SLAs for 25,000 workstations, 8,000 servers, and 600 internet-facing systems while limiting emergency changes to 20 per month; what SLA model would you set?
system-design - 26
Authenticated scanning covers only 62% of 60,000 managed endpoints, and the target is 95% within 90 days without granting scanners domain-admin rights; how would you close the gap?
endpoints - 27
Two scanners disagree on 4,000 findings across 1,200 internet-facing applications; how would you reconcile results and keep false-positive disputes under 10% of remediation time?
conflict - 28
There are 900 overdue critical findings, and application owners request 180 exceptions because fixes would miss a quarterly release; how would you govern exceptions while cutting overdue exposure by 70%?
error-handling - 29
A legacy manufacturing estate has 2,400 hosts, only 4 maintenance windows per year, and 130 KEV findings; how would you reduce exploitable exposure by 80% before the next window in 45 days?
attacks - 30
A weekly vulnerability queue adds 12,000 findings but engineering can remediate only 3,000; how would you prevent backlog growth while keeping all KEV findings inside SLA?
backlogvulnerabilitiesdata-structures - 31
Leadership wants 4 vulnerability metrics for 75,000 assets, but raw critical counts fluctuate by 30% with every scan; which measures would you implement to show real risk reduction over 2 quarters?
vulnerabilitiesrisk-managementmonitoring - 32
You have 30 days of CrowdStrike process and network telemetry from 18,000 Windows endpoints; how would you design a hunt for T1059.001 PowerShell that returns no more than 200 candidates for review?
designendpointsconcurrency - 33
CrowdStrike stores 14 days of endpoint events for 22,000 Windows hosts; how would you design a hunt for ATT&CK T1003.001 LSASS Memory with fewer than 100 high-confidence results?
designendpointsmemory - 34
Zeek retains 45 days of 2 billion DNS records for 12,000 clients; how would you design a hunt for T1071.004 DNS that analysts can complete in 5 working days?
designdns - 35
Microsoft Sentinel holds 90 days of AzureActivity from 140 subscriptions; how would you hunt for T1098.003 Additional Cloud Roles while limiting review to 80 actor-role pairs?
cloud-security - 36
You have 21 days of EDR policy and process data from 30,000 endpoints; how would you hunt for ATT&CK T1562.001 Impair Defenses with fewer than 120 candidates despite legitimate maintenance?
endpointsconcurrency - 37
Splunk contains 30 days of Windows events from 16,000 hosts; how would you design a hunt for T1021.002 SMB/Windows Admin Shares that returns at most 150 lateral-movement candidates?
designsiem - 38
CrowdStrike Falcon is installed on 28,000 of 32,000 endpoints, but 9% of sensors are more than 2 versions behind; how would you reach 98% active coverage and 95% currency within 60 days?
runtime-securityendpointscoverage - 39
Cortex XDR protects 6,000 production servers and 4,000 developer endpoints, but one prevention policy must not cause more than 0.5% workload failures; how would you design policy tiers and rollout?
designendpoints - 40
A data center hosts 120 applications on 4 flat VLANs, and the target is to reduce reachable application paths by 85% without exceeding 2 hours of downtime per application; how would you design segmentation?
network-securitydesign - 41
Two thousand production servers currently have unrestricted internet egress, but 300 require vendor APIs; how would you reduce allowed destinations by 90% while keeping failed update jobs below 1%?
procurementapi - 42
Six thousand employees hold 18,000 cloud entitlements, and 35% have not been used in 90 days; how would you remove 50% of unused privilege without causing more than 2% access-related tickets?
cloud-security - 43
A company has 450 administrators across 2,800 servers, 70 network devices, and 40 databases; how would you design PAM to eliminate 80% of standing privilege while keeping emergency access under 15 minutes?
databasedesign - 44
A federal contractor must evidence NIST 800-53 Rev. 5 audit controls for 300 systems every quarter; how would you implement AU-2, AU-3, AU-6, and AU-11 so evidence preparation takes under 2 days?
governancesystem-design - 45
An ISO 27001 scope includes 45 business processes and 160 applications, but 25 Annex A controls rely on manual evidence; how would you automate 70% of evidence before the next audit in 6 months?
governanceconcurrency - 46
A SaaS company needs SOC 2 evidence for quarterly access reviews across 4,500 users and 85 applications; how would you design the control so exceptions close within 10 business days?
compliancesoc-operationsdesign - 47
A healthcare SaaS with 70 services must satisfy FedRAMP Moderate and HIPAA for the same environment; how would you design shared access and audit evidence while limiting framework-specific work to 20%?
compliancedesigncloud - 48
A payment platform estimates a credential-abuse event could cost $2 million to $6 million and occur once every 4 to 10 years; how would you calculate ALE and decide whether a $350,000 annual control is justified?
estimation - 49
Using FAIR, a ransomware scenario affects 8 critical services, with threat-event frequency estimated at 0.2 to 1.0 per year and loss magnitude at $3 million to $18 million; how would you quantify the range for a $1.2 million segmentation proposal?
malwarenetwork-securityestimation - 50
A risk committee reviews 140 cyber risks quarterly and can discuss only 15; how would you set quantitative acceptance thresholds so high-tail risks are not hidden by low expected loss?
risk-management - 51
CrowdStrike detects a Cobalt Strike beacon on 14 Windows hosts, with SMB logons between 6 of them during the last 40 minutes; how would you contain and scope the incident without destroying evidence?
incidentsedr - 52
A domain admin account performs Kerberos service-ticket requests for 420 SPNs, then logs on to 9 servers from a workstation it has never used; what would you do in the first hour?
identity-access - 53
GuardDuty reports an AWS access key calling AssumeRole in 18 accounts and opening HTTPS sessions to a new VPS provider; how would you contain the cloud breach while preserving production workloads?
sessionscloud-securityhttps - 54
Falco detects a reverse shell from one Kubernetes pod, and audit logs show the service account listing secrets in 7 namespaces; how would you respond to the suspected cluster intrusion?
secretsruntime-securitynetwork-security - 55
Zeek shows one finance server sending 900 TXT queries per minute with 55-character labels to a 3-day-old domain, and EDR shows no known malware; how would you handle the suspected DNS C2?
malwareendpointsqueries - 56
An employee VPN session is followed by RDP connections to 11 servers and a 6 GB transfer over SMB in 25 minutes; the employee says they are asleep overseas, so how would you respond?
sessionsnetwork-security - 57
During an APT investigation, Sysmon records PsExec and WMI execution from an SCCM server to 32 endpoints, but that server also performs legitimate administration; how would you separate compromise from normal activity?
incident-responseendpoints - 58
A public IIS server shows a web shell, outbound C2, and a successful 4624 type 3 logon to a file server 12 minutes later; how would you scope a possible domain breach?
- 59
An implant on 5 developer laptops uses GitHub issue comments as C2 every 90 seconds, and blocking all GitHub would stop 600 engineers; what containment decision would you make?
incident-response - 60
EDR reports credential dumping on a domain controller, and a new scheduled task appears on 3 other domain controllers; how would you handle the possibility that the Active Directory forest is compromised?
incident-responseendpoints - 61
A new HTTPS C2 certificate is observed on 23 endpoints across 4 countries, but its IPs rotate every 6 minutes behind a CDN; how would you contain it without blocking the CDN?
network-securitycryptographyhttps - 62
A breach is discovered 21 days after the first beacon, with 2 TB of EDR data but only 7 days of full packet capture; how would you reconstruct scope and state your confidence?
endpoints - 63
Ransomware begins encrypting 420 of 9,000 Windows endpoints at roughly 35 hosts per minute; what containment sequence would you execute?
encryptionmalwareincident-response - 64
Ransomware has encrypted 18 of 35 ESXi hosts and vCenter is still reachable; how would you contain the virtual estate without powering off unaffected critical systems?
encryptionmalwaresystem-design - 65
A hospital sees ransomware on 160 of 1,800 endpoints, including 22 clinical workstations, and patient-care leadership resists network isolation; what would you decide?
malwareendpoints - 66
An attacker uses a compromised AWS role to copy 14 TB from S3 and starts encrypting objects with customer-managed KMS keys; how would you stop the cloud extortion event?
encryptionincident-responsecloud-security - 67
A destructive wiper presents a ransom note on 75 Linux servers, but disk changes do not match encryption; how would that change your response?
encryptioncryptography - 68
Ransomware compromises the backup console and deletes 60% of online restore points before endpoint encryption starts; what would you do to protect remaining recovery options?
encryptionmalwarecryptography - 69
Encryption stops after 96 hosts, but the attacker still has an active chat channel and claims to have stolen 800 GB; when would you move from containment to recovery?
encryptioncryptographyincident-response - 70
A phishing campaign targets 2,400 employees, 170 open the attachment, 17 submit credentials, and 4 accounts show successful sign-ins; how would you contain it?
social-engineering - 71
Thirty users approve a malicious OAuth consent grant requesting Mail.Read and Files.Read.All, but no passwords were stolen; how would you investigate and remediate it?
oauthpasswordsidentity-access - 72
A finance executive's mailbox sends changed bank details to 12 suppliers, and one supplier has already transferred $180,000; what are your first actions?
- 73
A QR-code phishing campaign reaches 8,000 mobile users, and 63 enter credentials through personal phones that lack EDR; how would you scope compromised accounts?
social-engineeringincident-responseendpoints - 74
An adversary-in-the-middle phishing kit steals session cookies from 9 users, and MFA resets do not stop new sign-ins; how would you contain the accounts?
sessionscookiessocial-engineering - 75
A compromised shared support mailbox has 46 delegates, and malicious replies were sent to 320 customers; how would you identify the responsible identity and contain the impact?
delegationincident-response - 76
A live phishing campaign changes sender domains every 20 minutes and has delivered 35 variants in 3 hours; how would you stop it without blocking legitimate partner mail?
social-engineering - 77
DLP reports an engineer copying 38 GB of design files to USB during the week before resignation; how would you investigate without tipping off the employee?
incident-responsedesign - 78
A data scientist downloads 120 GB from Snowflake to a personal cloud-storage domain over 2 nights, claiming it was model training; how would you determine whether this is insider exfiltration?
data-exfiltrationcloud-securitysnowflake - 79
A contractor clones 240 private Git repositories and uploads a 9 GB archive to an unknown VPS on the final contract day; how would you scope and preserve the case?
git - 80
A database administrator exports 65 GB from a customer database after receiving a poor performance review, but the DBA is still needed for a production migration in 6 hours; what would you do?
databasemigrationsperformance - 81
DLP flags a systems engineer transferring 400 GB to an external host, but the engineer says it is an approved backup test; how would you avoid both a missed insider incident and a damaging false accusation?
system-designbackupsincidents - 82
The SOC receives 14,000 alerts during a 6-hour outage, and one confirmed C2 alert is mixed into the backlog; how would you triage with 8 analysts?
backlogalertingsoc-operations - 83
Three analysts receive 22 simultaneous high-severity alerts across identity, endpoint, and cloud, but can actively investigate only 6; how would you choose?
incident-responsecloud-securityendpoints - 84
After a 4-hour SIEM outage, 9 million delayed events arrive and create 6,500 alerts whose event times differ from ingest times; how would you prevent bad triage decisions?
alertingsiem - 85
A vendor reports active exploitation of its VPN appliance, and the SOC suddenly sees 1,100 related alerts across 74 devices; how would you triage before a patch is available?
attacksnetwork-securitysoc-operations - 86
During alert overload, a manager asks you to suppress all medium-severity endpoint alerts for 24 hours; what safer decision would you make?
severity-priorityendpointsalerting - 87
A malware outbreak affects 3 factories in different time zones, and local IT teams begin taking conflicting containment actions; how would you coordinate incident response?
malwareincident-responseincidents - 88
An incident exposes records for an estimated 80,000 customers across 5 countries, but forensics can confirm only a range of 60,000 to 110,000; how would you coordinate technical, legal, and communications work?
estimationincidentszero-to-one - 89
A managed-service provider with access to 260 servers reports its remote-management platform is compromised; how would you coordinate containment with the vendor?
incident-responseprocurement - 90
A P1 incident spans teams in Singapore, London, and New York for 30 hours, and two handoffs have already lost key evidence; how would you fix coordination?
incidents - 91
Executives want a revenue-critical service reconnected 45 minutes after containment, but two persistence mechanisms remain unexplained; how would you make the decision?
incident-response - 92
A threat hunt finds one malicious SHA-256 on 27 endpoints, but only 4 executed it; what actions would you take?
threat-huntingendpoints - 93
A hunt discovers the same rare scheduled task on 8 servers, launching rundll32 every 17 minutes from a writable directory; how would you validate and contain it?
validation - 94
An identity hunt finds one FastPass device enrolling 14 new user accounts in Okta over 3 hours; what would you investigate and contain?
incident-response - 95
A hunt identifies a rare TLS certificate fingerprint communicating from 19 Linux servers, but the certificate also appears on 3 approved monitoring nodes; how would you avoid a bad blanket block?
tlscryptographymonitoring - 96
A hunt finds that 11 endpoints contacted a domain now sinkholed by law enforcement, but the last contact was 18 days ago; how would you use the stale indicator?
endpoints - 97
A junior analyst isolates 80 production servers after misreading a low-confidence IOC, causing a 25-minute outage; how would you coach them and repair the process?
concurrency - 98
A capable SOC analyst misses 3 escalations in one month because their cases lack timelines and evidence links; how would you mentor them?
escalationmentoringsoc-operations - 99
Your 12-person SOC uses five different investigation styles, and handoff rework is 28%; how would you mentor the team toward consistent quality without forcing scripts?
incident-responsesoc-operationsmentoring - 100
Two new SOC analysts must be ready to handle a ransomware P1 within 30 days, but production access cannot be used for training; how would you mentor and assess them?
malwaresoc-operationsmentoring