Questions d'entretien : Security Engineer
100 vraies questions avec réponses modèles et explications pour les candidats Ingénieur senior en sécurité.
Voir un exemple de CV : Security Engineer →S'entraîner avec des cartes
Répétition espacée · Hunter Pass
Questions
I would model the control plane and data plane as separate DFD trust zones, then make tenant identity an invariant on every STRIDE path.
- The DFD names 5 boundaries: CDN, API gateway, control services, tenant workers, and PostgreSQL, with tenant_id carried in a signed internal token.
- I apply STRIDE to each flow and prioritize cross-tenant reads, confused-deputy jobs, and control-plane takeover above single-tenant denial of service.
- I validate the model with 12 abuse cases, including forged tenant_id, stale job claims, and a control service calling a tenant worker without audience binding.
- The design is accepted only when isolation tests cover 100% of data APIs and the control plane preserves the 99.99% SLA during one tenant's rate-limit event.
Pourquoi cette question est posée: The interviewer is testing whether the candidate can turn a DFD and STRIDE into enforceable tenant-isolation invariants at real SaaS scale.
I would keep primary account numbers inside Stripe-hosted collection and place only 3 token-aware services in the cardholder-data trust zone.
- Stripe Elements posts card data directly to Stripe, while the checkout API receives a PaymentMethod token and never logs request bodies on that route.
- The payment orchestrator, webhook verifier, and immutable ledger form the scoped zone; mTLS and separate Kubernetes namespaces prevent 140 other services from calling ledger write endpoints.
- A DFD marks browser-to-Stripe, Stripe-to-webhook, and orchestrator-to-ledger boundaries, with idempotency_key and Stripe-Signature verified before state changes.
- Quarterly PAN canary scans and 100% egress allowlisting prove that no card data enters Kafka, Datadog, or the general data lake.
Pourquoi cette question est posée: A strong answer confines payment data with explicit artifacts instead of claiming that tokenization alone removes PCI DSS scope.
I would separate upload acceptance from content trust by placing every S3 object in a quarantine state until the scanning pipeline attests it.
- The API issues a 15-minute presigned URL bound to tenant, object key, size, and checksum, so 20 GB never crosses the application tier.
- An S3 event sends the version ID to an isolated ClamAV and YARA scanner with no production credentials, and the worker writes a signed ScanResult attestation.
- A promotion service copies only clean versions into a read bucket; all download paths require the promoted object tag and cannot address quarantine keys.
- The 300 ms receipt means accepted, not trusted, while a 30-minute scan SLO and a 24-hour quarantine TTL bound backlog and storage cost.
Pourquoi cette question est posée: The interviewer is evaluating whether large-upload performance is separated cleanly from the decision to trust content.
I would authenticate at the webhook edge, preserve the raw event, and make replay handling provider-specific through a durable WebhookEnvelope.
- The edge verifies each named scheme such as Stripe-Signature or GitHub X-Hub-Signature-256 against the raw body before JSON parsing.
- WebhookEnvelope stores provider, key ID, event ID, timestamp, body hash, and verification result in Kafka with 7-day retention.
- A Redis SETNX key on provider plus event ID blocks duplicates for 72 hours, while the database uses the same pair as a unique idempotency constraint.
- Providers without event IDs receive a 5-minute timestamp window and HMAC body-hash key, with a documented exception for 2 legacy providers behind mTLS.
Pourquoi cette question est posée: A strong answer distinguishes cryptographic authenticity, bounded replay prevention, and idempotent business processing.
I would treat every logistics firm as a separate external trust domain and terminate its identity at an Envoy API gateway before order services.
- Each partner gets a dedicated OAuth 2.0 client, mTLS certificate, audience, and 500-request-per-second quota rather than a shared partner role.
- The DFD marks internet, gateway, authorization service, order API, and audit sink boundaries; internal calls carry partner_id and customer account as signed claims.
- Abuse cases cover BOLA on order IDs, bulk enumeration, stale client credentials, schema over-posting, and a partner using another firm's shipment callback.
- OPA policy tests include 40 negative fixtures, and the gateway caches allow decisions for 30 seconds to preserve the 200 ms p99 target.
Pourquoi cette question est posée: The interviewer is checking whether partner access is isolated per organization and tested against concrete API abuse paths.
I would make impersonation a separate high-trust workflow, not another permission inside the normal support session.
- The console requests a 15-minute ImpersonationGrant from a dedicated service after WebAuthn step-up, a ticket ID, and approval by 1 of 25 authorized agents.
- The grant names agent, tenant, user, allowed actions, and reason; the PEP rejects export, billing changes, credential reset, and cross-tenant navigation.
- UI and API responses carry an impersonation banner and session ID, while an append-only AuditEvent records every read and mutation for 7 years.
- STRIDE review focuses on grant spoofing, approval tampering, tenant-context loss, and repudiation, with 20 negative authorization tests in CI.
Pourquoi cette question est posée: A strong answer turns support impersonation into a narrowly scoped, independently audited security boundary.
I would make tenant provenance immutable from Kafka ingestion through Trino authorization and prevent user-supplied data from defining that provenance.
- Kafka topics carry a broker-authenticated tenant header, and Spark rejects records whose payload tenant_id disagrees with the signed ingestion context.
- Iceberg tables partition by tenant bucket but rely on Ranger row filters and per-tenant encryption context, because partitions alone are not an access boundary.
- Trino receives tenant claims from OIDC and applies Ranger policy at query planning; result caches include tenant and policy version in the key.
- The threat model tests 25 paths including forged headers, Spark joins across tenants, stale Ranger cache, and exported query results, while preserving 10,000 queries per minute.
Pourquoi cette question est posée: The interviewer is evaluating end-to-end tenant provenance rather than a single database permission.
I would treat producing, transporting, interpreting, and consuming an event as 4 distinct trust decisions.
- SPIFFE identities and Kafka ACLs bind each of 300 services to named topics, with no wildcard write access and separate dead-letter permissions.
- Apicurio Registry accepts only signed OrderCreated schema versions, and consumers reject unknown major versions rather than deserializing arbitrary fields.
- Each event includes producer identity, tenant, event ID, schema digest, and creation time in a signed EventEnvelope; consumers still authorize the requested business action.
- Chaos tests revoke 1 producer and corrupt 1% of envelopes to prove bad events quarantine without reducing the 99.95% order SLA.
Pourquoi cette question est posée: A strong answer does not confuse broker authentication with trust in event meaning or consumer authorization.
I would make the GraphQL gateway a policy enforcement boundary for operation shape, identity, and downstream data ownership.
- Apollo Router accepts persisted queries only, caps depth at 12 and computed cost at 5,000, and rejects introspection outside 2 non-production environments.
- The schema marks each field with owning service and required scope, while OPA checks tenant and object relationship before a resolver fetches data.
- DataLoader caches are request-scoped and keyed by tenant plus object ID, preventing one of 25 clients from receiving another tenant's cached result.
- Red-team fixtures cover aliases, fragments, batching, nested BOLA, and cost evasion, and policy evaluation gets 10 ms of the 250 ms p99 budget.
Pourquoi cette question est posée: The interviewer is testing whether GraphQL-specific composition risks become measurable gateway and resolver controls.
I would make residency a cryptographically bound routing attribute and fail over compute without silently moving protected records across regions.
- The TenantResidency record is signed by the control plane and maps each tenant to eu-central-1 or us-east-1 before any storage key is issued.
- Regional API, PostgreSQL, S3, KMS, and audit stacks have no cross-region data replication; only schema and non-sensitive configuration replicate globally.
- Global DNS can fail EU traffic to a second Frankfurt availability zone within 5 minutes, but cannot route it to Virginia even if that lowers availability.
- A DFD and 16 failover tests verify backup, support, analytics, and log-export paths against the 12-year residency rule and 99.99% SLA.
Pourquoi cette question est posée: A strong answer treats residency as a hard trust boundary with an explicit availability trade-off.
I would centralize workforce identity in Entra ID and issue application-specific OIDC sessions rather than a reusable enterprise bearer token.
- WebAuthn passkeys are mandatory for 3,000 privileged users, while the remaining workforce migrates within 6 months from push MFA.
- Each of 90 applications validates issuer, audience, nonce, PKCE, signature, and a 10-minute ID token lifetime through a maintained OIDC middleware.
- SCIM disables application accounts within the 4-hour SLA, and high-risk apps also consume continuous access evaluation events for revocation under 15 minutes.
- Two hardware-key break-glass accounts bypass federation only through a recorded 30-minute procedure and are tested quarterly.
Pourquoi cette question est posée: The interviewer is evaluating token validation, lifecycle deprovisioning, and bounded recovery rather than OIDC login alone.
I would bind every federation connection to an immutable tenant identifier and never select a tenant from an email domain or unverified claim.
- A login starts from a tenant-specific URL or signed discovery code, then maps the exact OIDC issuer plus subject or SAML entityID to one Connection record.
- The callback validates state, nonce, PKCE, audience, ACS URL, and signed RelayState before creating a session containing tenant_id and connection_id.
- JIT provisioning is limited to approved group claims, while SCIM remains authoritative for 1,200 enterprise tenants that require deprovisioning under 2 hours.
- A 30-case conformance suite tests shared domains, duplicate subjects across issuers, IdP-initiated SAML, and connection deletion.
Pourquoi cette question est posée: A strong answer makes federation connection identity, not a mutable email attribute, the tenant boundary.
I would issue short-lived SPIFFE identities from regional SPIRE trust domains and derive authorization from workload attributes rather than secrets.
- SPIRE Server uses Kubernetes PSAT node attestation, and agents attest namespace plus service account before issuing a 30-minute X.509-SVID.
- Each of 40 clusters has an intermediate signing authority, while 4 regional servers chain to an offline root and federate only named partner trust domains.
- Envoy validates SVID trust domain and SPIFFE ID, then OPA authorizes exact source-to-destination methods instead of accepting any valid certificate.
- Rotation occurs before 50% of lifetime, with a 99.99% issuance SLO and cached SVIDs covering a 15-minute SPIRE outage.
Pourquoi cette question est posée: The interviewer is testing the full workload-identity lifecycle from attestation through service authorization and failure behavior.
I would use RBAC for stable job capabilities and ABAC only for contextual narrowing, with Cedar as the single policy language.
- About 120 business roles replace 2,000 job titles, and each role grants named actions such as payment.read or payment.approve.
- Cedar conditions restrict those grants by tenant, region, data classification, device trust, and transaction amount from the 40 available attributes.
- Separation of duties prevents one identity from creating and approving payments above $10,000, and no ABAC rule can expand a missing RBAC grant.
- Policy CI runs 5,000 allow and deny fixtures plus differential tests against the previous bundle before a signed policy version reaches PDPs.
Pourquoi cette question est posée: A strong answer controls role explosion while keeping contextual policy monotonic and testable.
I would model document access as versioned relationship tuples and evaluate it in a Zanzibar-style authorization service separate from document storage.
- Tuples such as team:42#member@user:7 and document:9#parent@folder:3 support nested membership without copying ACLs onto 800 million documents.
- Share links become expiring link principals with viewer only, a 7-day maximum TTL, and no right to create further relations.
- The PEP sends object, relation, subject, tenant, and consistency token; the PDP returns allow plus the tuple revision used for audit.
- A 10 ms local cache serves unchanged checks, but writes use read-your-writes tokens so revocation remains under 60 seconds and total p99 stays below 50 ms.
Pourquoi cette question est posée: The interviewer is evaluating graph authorization semantics, consistency, and cache safety at large object counts.
I would put PEPs in every Envoy sidecar and run replicated OPA PDPs per cluster, with signed policy bundles and bounded local caching.
- Envoy ext_authz sends caller SPIFFE ID, method, route template, tenant, and resource attributes to an OPA instance in the same node pool.
- OPA receives a signed bundle every 60 seconds from a central policy control plane, but decisions use local data so 120,000 checks per second do not cross regions.
- Allow cache entries live 5 seconds and include policy revision plus subject, action, resource, and tenant; explicit denies are never replaced by stale allows after revocation events.
- The PEP fails closed for mutations and permits a 2-minute last-known-good mode only for 20 named read routes, keeping policy p99 under 8 ms.
Pourquoi cette question est posée: A strong answer separates central policy distribution from low-latency local enforcement and defines failure semantics.
I would enforce tenant isolation with PostgreSQL RLS derived from authenticated database identity, then verify it with database-level negative tests.
- An authentication broker issues 15-minute credentials for exactly one tenant login role; the pool never reuses a connection across database roles and runs DISCARD ALL before reuse.
- RLS calls a fixed-search-path SECURITY DEFINER tenant_for_session function that maps immutable session_user through a protected table owned by a no-login policy role.
- Tenant roles cannot assume another role, alter the mapping, use BYPASSRLS, or own tables; FORCE ROW LEVEL SECURITY, USING and WITH CHECK apply to every tenant table, and foreign keys include tenant_id.
- Partitioning by tenant hash is only a scale control, while 10,000 randomized cross-tenant reads and writes plus production canaries verify zero rows outside the authenticated tenant.
Pourquoi cette question est posée: The interviewer is checking defense in depth at the database boundary, including pool reuse and relational integrity.
I would derive tenant scope from the authenticated session at both OpenSearch and Redis boundaries, never from a client filter.
- OpenSearch uses document-level security on tenant_id and separate index aliases for 200 regulated tenants, with the application role unable to query raw backing indices.
- Every Redis key starts with tenant ID, authorization policy version, user visibility class, and normalized query hash; a global cache key is prohibited by the cache SDK.
- Bulk indexing rejects documents whose signed ingestion tenant differs from tenant_id, and aliases cannot be changed by the ingest role.
- A 5,000-case isolation suite covers aggregations, suggesters, scroll IDs, cache invalidation, and alias access while sustaining 20,000 searches per second.
Pourquoi cette question est posée: A strong answer protects secondary stores and caches as first-class authorization boundaries.
I would use OAuth 2.0 client credentials with RFC 8705 mTLS certificate-bound access tokens for each partner workload.
- Each of 45 partners registers separate production and test certificates, and the authorization server places the x5t#S256 thumbprint plus treasury audience in a 5-minute JWT.
- The API gateway validates certificate chain, token signature, issuer, audience, expiry, and cnf thumbprint before forwarding partner_id.
- Scopes separate transfer.create, transfer.read, and beneficiary.manage, while amounts above $100,000 require a second signed approval assertion.
- Certificates rotate every 90 days with a 7-day overlap, and revoked thumbprints reach all gateways through a 60-second deny feed.
Pourquoi cette question est posée: The interviewer is evaluating proof-of-possession token binding, partner isolation, and practical credential rotation.
I would replace standing administrator roles with a just-in-time AccessGrant issued by a dedicated privilege service.
- The request names resource, 1 of 8 approved tasks, ticket, duration, and peer approver; WebAuthn step-up is required before a 30-minute grant.
- AccessGrant becomes a short-lived cloud role or Kubernetes credential with session tags, never a reusable group membership or static kubeconfig.
- PEPs consume a revocation stream and reject the grant within 60 seconds, while all kubectl and cloud API calls retain grant ID in audit logs for 2 years.
- Two break-glass roles require dual hardware keys, page the security on-call immediately, and are exercised every 90 days.
Pourquoi cette question est posée: A strong answer removes standing privilege while defining issuance, enforcement, revocation, and recovery artifacts.
Questions verrouillées
- 21
A storage service encrypts 8 PB across 20,000 tenants and writes 1 million objects per minute; how would you design envelope encryption with AWS KMS?
encryptioncryptographyenvelope-encryption - 22
A checkout fleet needs 70,000 decrypt operations per second, but AWS KMS allows 30,000 and payment p99 must remain under 180 ms; how would you redesign key usage safely?
- 23
A certificate authority signs 25 million device certificates per year, requires FIPS 140-3 Level 3 key protection, and must survive one data-center loss; how would you design the HSM hierarchy?
cryptographyhsmdesign - 24
A service mesh has 15,000 workloads in 30 clusters and needs certificate rotation every 24 hours with no more than 0.01% failed handshakes; how would you design PKI and mTLS rotation?
mtlscryptographydesign - 25
A data platform must rotate a compromised KMS key protecting 6 billion records within 72 hours without taking reads offline; what cryptographic migration design would you choose?
migrationsdesignincident-response - 26
A webhook platform signs 50,000 events per second for 1,200 customers, and customers may take 14 days to rotate Ed25519 keys published through JWKS; how would you design the signing protocol and key lifecycle?
designwebhookstyping - 27
An audit service ingests 4 million security events per minute and must prove 7-year integrity using S3 Object Lock and Merkle checkpoints; how would you design tamper-evident storage?
designaws - 28
A privacy service must use tenant-specific KMS keys to make 900 TB of backups cryptographically unrecoverable within 24 hours while retaining shared backup media for 5 years; how would you design crypto-shredding?
designbackups - 29
A global messaging service encrypts 2 million messages per second with AWS KMS in 6 regions and must survive a region loss with RTO 10 minutes; how would you design the KeySet architecture?
designencryption - 30
An identity service stores credentials for 35 million users and must verify 8,000 logins per second with a 250 ms p99 target; how would you design password and recovery-token cryptography?
passwordstokensdesign - 31
An engineering organization has 850 repositories, 4 languages, and 12,000 pull requests per month; how would you design a GitHub Actions secure SDLC platform without copying CI logic into every repo?
code-reviewdesignci-cd - 32
A portfolio of 500 web services produces 70,000 weekly SAST and DAST findings; how would you correlate Semgrep, CodeQL, ZAP, and IAST results into actionable defects?
sastdastiast - 33
A company ships 300 Java services daily and receives 18,000 SCA findings per week; how would you prioritize dependency risk using reachability, EPSS, and runtime context?
dependenciesrisk-management - 34
A retail platform has 220 staging services and releases every 30 minutes; how would you deploy IAST without adding more than 8% latency or flooding developers?
deploymentiastlatency - 35
A bank has 600 repositories and a 10-minute CI budget; what risk-based CodeQL gates and ReleaseAttestation checks would you enforce for pull requests and production releases?
code-reviewrisk-management - 36
A security platform serves 1,100 repositories and allows OPA PolicyException waivers, but no exception may outlive 30 days; how would you design exception governance as code?
policydesignerror-handling - 37
A Bazel monorepo contains 1,800 deployable components and 9 million lines of code, but changed-file detection misses transitive risk; how would you scope security scans under a 15-minute p95 target?
detectionrisk-managementdeployment - 38
A security program covers 900 repositories and claims 95% SAST coverage, but leadership needs outcome metrics in a dbt semantic layer for the next 2 quarters; what measurement architecture would you build?
monitoringdbtsast - 39
An architecture review team handles 250 SecurityChangeManifest submissions per month with a 5-day SLA; how would you automate DFD and STRIDE triage without replacing deep review?
stridearchitecture - 40
Forty product teams use Backstage to create 300 new services per year and need a secure production path in under 1 day; how would you design golden templates and self-service controls?
design - 41
A build platform produces 10,000 container images per day from 700 repositories; how would you reach SLSA Build Level 3 style provenance without trusting repository runners?
supply-chaincontainers - 42
An enterprise inventories 6 million components across 25,000 artifacts and must answer a new CVE exposure query within 30 minutes; how would you architect SBOM ingestion and lookup?
supply-chainvuln-managementcve - 43
A medical-device vendor has 1,500 products and must publish VEX for 9,000 CVE matches within 48 hours; how would you design defensible VEX generation?
vuln-managementcvedesign - 44
A registry stores 4 million artifacts and 60 Kubernetes clusters deploy Cosign-signed images 30,000 times per day; how would you enforce artifact signing without a shared private key?
kubernetesdeploymentartifacts - 45
A platform rebuilds 500 Cosign-signed base images with CycloneDX SBOMs weekly for 3,000 services and requires critical fixes within 24 hours; how would you design the trusted supply chain?
supply-chaindesign - 46
A Kubernetes platform runs 18,000 pods in 45 clusters and must reject privileged or unverified workloads with admission p99 below 100 ms; what policy architecture would you use?
kubernetesarchitecture - 47
A Kubernetes estate runs 1,600 SPIFFE-identified services across 45 clusters, generates 900,000 outbound requests per second, and permits only 120 external destinations with at most 3 ms p99 egress overhead; how would you segment and control workload egress?
kubernetes - 48
A vulnerability platform tracks 2 million AssetIDs and 500,000 CVE and CWE findings with remediation SLAs from 24 hours to 90 days; how would you design its risk and ownership architecture?
vulnerabilitiesvuln-managementrisk-management - 49
A build platform serves 900 repositories and must eliminate static cloud keys while allowing 20-minute deployments to 120 accounts; how would you design build identity and secret access?
secretscloud-securitydesign - 50
A regulated compiler pipeline builds 2,000 release artifacts per day and requires bit-for-bit reproducibility from 3 independent builders; how would you design hermetic builds and verification?
ci-cdartifactsdesign - 51
Credential-stuffing attempts jump from 800 to 24,000 per minute, successful logins from new devices triple, and support reports 37 takeovers; do you force-reset all 2 million accounts or contain selectively, and what is your decision?
- 52
A production JWT signing key appears in a public repository for 46 minutes, access tokens live for 60 minutes, and 180 services trust it; do you rotate immediately despite a likely 8-minute login outage, what is your decision, and what evidence defines recovery?
jwttokens - 53
An authorization-policy release makes 0.7% of 90,000 requests per minute return data outside the caller's role for 11 minutes; do you roll back the whole release or patch the rule in place, and what is your decision?
authidentity-accessrollback - 54
A shared-schema SaaS query omits tenant_id and returns 2,430 invoice rows from 19 tenants to 6 users over 23 minutes; do you take all billing reads offline or isolate one query path, and what is your decision?
queriesschemacloud - 55
An OAuth client with 480,000 monthly users accepts a wildcard redirect URI, and telemetry shows 312 authorization codes sent to a newly registered subdomain in 2 hours; do you disable the client or keep login available, and what is your decision?
authoauthidentity-access - 56
After an OIDC provider rotates keys, 14 of 180 services accept attacker-forged HS256 tokens carrying an unknown kid for 9 minutes because their verifier falls back to the first RSA JWKS key and uses its public-key bytes as an HMAC secret; do you block all federated login or quarantine those services, and what is your decision?
secretstokens - 57
Refresh-token reuse detection fires for 860 token families across 11 countries within 30 minutes, while 28% of mobile clients retry after network loss; do you revoke every family or suppress the detector, and what is your decision?
tokensdetectiontoken-reuse-detection - 58
A recovery endpoint bypasses MFA when three knowledge questions are answered, and 74 of 20,000 daily recoveries came from one proxy network with 19 confirmed takeovers; do you shut down recovery for everyone, and what is your decision?
identity-accessendpointsproxy - 59
A service-account bearer token was printed in centralized logs viewed by 260 engineers for 6 days, and it can write to 4 production queues; do you rotate only the token or also pause its consumers, and what is your decision?
tokensdata-structures - 60
An administrator removed a user's privileged role, but a 30-minute authorization cache allowed 17 destructive actions over 12 minutes; do you disable caching globally or add targeted revocation, and what is your decision?
authidentity-accesscaching - 61
A GitHub Actions job exposed a production database secret in public logs for 18 minutes, 43 clones occurred, and the credential grants writes to 12 schemas; do you stop all deployments or rotate in place, and what is your decision?
databaseschemaci-cd - 62
Version 4.8.2 of a payment dependency begins sending DNS queries to an unknown domain from 9 of 140 pods, and 62% of checkouts use that version; do you roll back immediately or wait for vendor confirmation, and what is your decision?
dependenciesrollbackqueries - 63
A typosquatted npm package ran a postinstall script on 27 CI runners, and 6 runners held cloud credentials valid for 45 minutes; do you rebuild every artifact from that day or only the 83 that imported the package, and what is your decision?
npmartifactscloud-security - 64
A maintainer account for a library used by 320 services is compromised, 2 malicious releases were available for 3 hours, and 71 builds resolved them; do you freeze all releases or allow unaffected teams to continue, and what is your decision?
incident-response - 65
Cosign verification fails for 186 of 4,000 images after a registry migration, but 22 unsigned images are already running in production; do you evict them immediately or preserve availability, and what is your decision?
migrationsregistries - 66
SLSA provenance for 54 release artifacts names the correct repository but an unexpected builder identity, and 8 artifacts reached 15% of users; do you roll back all 54 or investigate first, and what is your decision?
supply-chainincident-responserollback - 67
The registry tag payments-api:2026.07.14 moved from the approved digest ending 91af to a different, validly signed diagnostic-image digest ending c42e that exposes port 9001; 17 clusters pulled it and 6 ran it for 23 minutes; do you trust the signature or roll back every tag consumer, and what is your decision?
rollbackregistriesapi - 68
A shared CI cache returns a binary whose SHA-256 differs across 3 of 12 runners, and 240 builds consumed the cache in 7 hours; do you purge the cache and rebuild everything or isolate only the 3 runners, and what is your decision?
caching - 69
A third-party GitHub Action tag used by 600 repositories was force-moved for 35 minutes, and 114 workflows ran it; do you disable GitHub Actions company-wide or block that action and rebuild, and what is your decision?
ci-cddependencies - 70
An artifact-signing service issues 73 signatures outside its normal release window, 11 signatures cover unknown digests, and the key sits in an HSM; do you disable signing and delay 6 scheduled releases, and what is your decision?
hsmartifacts - 71
Falco reports a shell and cryptominer in 18 of 2,800 product pods, node CPU reaches 96%, and checkout p99 doubles to 900 ms; do you cordon all 6 nodes or kill only the pods, and what is your decision?
runtime-security - 72
The Kubernetes admission webhook times out for 14 minutes, fail-closed blocks 320 deployments, and 7 emergency fixes await release; do you switch it to fail-open, and what is your decision?
kubernetesdeploymentwebhooks - 73
Kubernetes audit logs show a privileged pod mounting /var/run/containerd.sock for 6 minutes in 1 of 24 clusters, and it queried 430 secrets; do you shut down the cluster or isolate the node and namespace, and what is your decision?
kubernetescontainerssecrets - 74
Terraform drift opens an admin API to 0.0.0.0/0 for 51 minutes, 8,400 requests arrive from 73 IPs, and 16 return HTTP 200; do you roll back the whole stack or patch the exposure first, and what is your decision?
terraformrollbackiac - 75
A new WAF rule blocks an active SQL injection payload but also rejects 6.8% of checkout requests and costs $180,000 per hour; do you disable the rule or accept the loss, and what is your decision?
injectionweb-attackssql - 76
An API client enumerates sequential order IDs at 900 requests per second and receives 12,600 orders belonging to 340 users in 17 minutes; do you take the orders API offline or block the client and affected route, and what is your decision?
network-securityapi - 77
A PDF preview service makes 3,200 requests to the cloud metadata IP in 9 minutes, and 41 responses return credentials valid for 60 minutes; do you disable all uploads or only previews, and what is your decision?
cloud-security - 78
A GraphQL query with 18 nested fragments drives database CPU to 94% and p99 to 11 seconds across 70% of users; do you disable GraphQL or reject the query shape, and what is your decision?
databasequeriesgraphql - 79
A critical container-runtime CVE has public exploitation, 46 of 900 product nodes are vulnerable, and patching each node removes 8 minutes of capacity; do you drain immediately or wait for the weekend, and what is your decision?
vulnerabilitiesattacksvuln-management - 80
An eBPF runtime sensor sees exec from a writable /tmp in 32 API pods across 4 namespaces, but the sensor has a 0.4% false-positive rate; do you isolate all 32 pods or observe longer, and what is your decision?
kubernetesapi - 81
KMS decrypt calls rise from 4,000 to 85,000 per minute, 92% use one role, direct KMS cost is about $15.30 per hour at $0.03 per 10,000 requests, and throttling causes 12% checkout failures worth $14,000 in revenue per hour; do you disable that role, and what is your decision?
- 82
During an HSM key ceremony, 1 of 3 custodians cannot attend, the current signing key expires in 6 hours, the normal process requires all 3, and policy documents a witnessed break-glass 2-of-3 ceremony; do you invoke break-glass or delay releases, and what is your decision?
hsmconcurrency - 83
A TLS private key for api.example.com is found in a support bundle downloaded by 14 customers over 3 days, and 38% of clients cannot update pins quickly but were preprovisioned with a backup SPKI pin; do you revoke the certificate immediately, and what is your decision?
tlscryptographyapi - 84
A certificate intermediate expires in 42 minutes, 1,600 services use it, and rotation tests show 3.5% of Java clients reject the new chain; do you deploy the new chain or extend the old one, and what is your decision?
cryptographydeploymenttesting - 85
A review finds 680 million customer records encrypted with 3DES, regulators require a plan in 72 hours, and re-encryption would consume 40% database capacity for 18 days; do you block writes or migrate online, and what is your decision?
encryptioncryptographydatabase - 86
Telemetry finds 1.2 million AES-GCM ciphertexts sharing 18,000 nonces under the same key after a counter reset, affecting 46 tenants; do you take encrypted messaging offline, and what is your decision?
encryptioncryptography - 87
A logging change writes full access tokens and 240,000 customer email addresses to Splunk for 9 hours, with 1,100 employees able to search the index; do you delete the index or preserve it for investigation, and what is your decision?
tokensincident-responselogging - 88
A backup export containing 78 million customer records was publicly readable for 26 minutes, CDN logs show 9 complete downloads, and object encryption remained enabled; do you classify this as a confirmed breach, and what is your decision?
encryptioncryptographybackups - 89
Audit ingestion has a 47-minute gap during which 23 privileged database changes occurred, and two administrators deny making them; do you restore from backup or keep production running, and what is your decision?
databasebackups - 90
After a KMS rotation, 2.4% of 36 million documents fail to decrypt and customer support receives 620 tickets in 20 minutes; do you roll back the key alias or continue migration, and what is your decision?
migrationsrollback - 91
A researcher reports unauthenticated remote code execution with a working exploit against version 7.4, used by 38% of customers, and requests acknowledgment within 24 hours; do you validate first or immediately disable the feature, and what is your decision?
validationattacks - 92
A critical SSRF report affects 420 enterprise tenants, the reporter plans publication in 7 days, and a full parser fix needs 14 days; do you request delay or ship a partial mitigation, and what is your decision?
ssrf - 93
A deserialization vulnerability is actively exploited at 27 requests per minute, the vendor patch arrives in 5 days, and the API handles 18,000 requests per second; do you disable the vulnerable operation or rely on a virtual patch, and what is your decision?
vulnerabilitiesattacksvulnerability-management - 94
A critical library patch removes an actively exploited CVE but raises checkout errors from 0.2% to 4.6% in a 10% canary; do you roll back the patch or expand it, and what is your decision?
attacksvuln-managementcve - 95
A researcher demonstrates cross-tenant file access affecting 64 of 8,000 tenants and plans a conference talk in 10 days; legal asks you to dispute severity, so do you challenge the rating or accept critical and remediate, and what is your decision?
severity-priority - 96
At 02:00, 1,900 accounts show impossible travel, 14 engineers join the call, and executives request updates every 15 minutes; do you lead hands-on investigation or act as incident commander, and what is your decision?
incidentsjoinsincident-response - 97
A possible data-exfiltration alert has 35% confidence, covers 4 TB of outbound traffic, and could cost $2 million per hour if the pipeline is stopped; do you declare severity 1 and halt it, and what is your decision?
alertingci-cdseverity-priority - 98
A junior engineer proposes disabling webhook signature verification to recover 28% failed deliveries before a launch in 3 hours; do you approve a 30-minute exception, and what is your decision?
webhookserror-handling - 99
During final review 90 minutes before release, you find one bulk-export endpoint lacks tenant authorization, while 12 enterprise customers expect the feature today; do you block the release or accept a temporary WAF rule, and what is your decision?
authidentity-accessendpoints - 100
A teammate merges a secret into a repository despite 2 approvals, the secret is exposed for 22 minutes, and this is the third review escape in a quarter; do you focus on the individual or change the review system, and what is your decision?
secretssystem-design