Skip to content

AWS Engineer interview questions

100 real questions with model answers and explanations for Senior Cloud Engineer candidates.

See a AWS Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

I would use Control Tower with workload OUs split by environment and regulatory boundary.

  • Keep Log Archive and Audit accounts in a Security OU, with delegated security services outside workload administration.
  • Create Infrastructure, Sandbox, Nonproduction, Production, and Regulated OUs so SCPs follow stable trust boundaries rather than team names.
  • Apply mandatory detective and preventive controls at registration, then add stricter data-region and network controls to the Regulated OU.
  • Move accounts between OUs only through a reviewed workflow because inherited SCP changes can immediately remove permissions.

Why interviewers ask this: The interviewer is evaluating whether the candidate can turn a 60-account organization into explicit administrative and compliance boundaries.

design

I would keep a small deny-based SCP set aligned to organization-wide invariants and OU-specific exceptions.

  • One baseline blocks leaving the organization, disabling security services, and altering central log destinations.
  • Separate SCPs enforce approved regions, regulated-service restrictions, and production network rules so each has one change owner.
  • Use IAM policies and permission boundaries for grants because SCPs define the maximum permission set and never grant access.
  • Test effective policies in a canary OU with representative roles before promoting them through the OU hierarchy.

Why interviewers ask this: The interviewer is checking precise knowledge of SCP semantics, quotas, and safe composition at organizational scale.

sessionsiamconfig

I would federate the corporate identity provider into IAM Identity Center and assign access through synchronized groups.

  • Define permission sets for job functions such as ReadOnly, Developer, Billing, and PlatformAdmin, with a 1-hour session duration.
  • Assign groups to account sets through automation and prohibit direct user assignments so joiner and leaver changes remain auditable.
  • Require phishing-resistant MFA at the identity provider and use short-lived role credentials instead of IAM access keys.
  • Maintain two tightly monitored break-glass roles outside federation, with hardware MFA and quarterly access tests.

Why interviewers ask this: The interviewer is testing whether identity design scales to 800 people while preserving short-lived access and an emergency path.

delegationconfigiam

I would reserve the Organizations management account for billing and organization changes, not daily security operations.

  • Delegate GuardDuty, Security Hub, and IAM Access Analyzer to a Security Tooling account owned by the security team.
  • Use a separate Audit account for read-only investigations and Config aggregation so evidence access is distinct from control changes.
  • Enable organization-wide auto-enrollment for new accounts and regions to prevent account vending from creating coverage gaps.
  • Monitor delegated-admin changes through organization trails and alert on any assignment outside the approved account IDs.

Why interviewers ask this: The interviewer is evaluating separation of duties and centralized service administration across 60 accounts.

I would expose an Account Factory for Terraform or Control Tower Account Factory workflow backed by a reviewed account request schema.

  • The request captures owner, cost center, environment, data class, network profile, and target OU before CreateAccount is called.
  • An asynchronous pipeline waits for account creation, applies baseline roles, logging, budgets, identity assignments, and network attachments.
  • Idempotency keys and a request state machine prevent duplicate accounts when Organizations APIs throttle or callers retry.
  • Measure request-to-ready p95 against the 30-minute target and pre-open support cases when organization account quotas approach 80 percent.

Why interviewers ask this: The interviewer wants a concrete self-service design that handles asynchronous AWS APIs, governance, and a 30-minute delivery objective.

computeaws

I would maintain quota requirements as part of the account and workload metadata rather than discover limits at launch.

  • Check Organizations account capacity, regional EC2 On-Demand vCPU quotas, EIP limits, NAT gateways, and any service-specific ceilings.
  • Submit quota increases from each target account through Service Quotas automation because many quotas are regional and account-scoped.
  • Run a capacity rehearsal with the intended instance families and Availability Zones at least 2 weeks before launch.
  • Where capacity is contractual, combine Capacity Reservations with an approved fallback region instead of relying on quota approval alone.

Why interviewers ask this: The interviewer is checking that the candidate distinguishes adjustable quotas from actual regional capacity and plans weeks ahead.

designgatewaynetworking

I would deploy one Transit Gateway per region in a central Network account and share it through Resource Access Manager.

  • Use separate TGW route tables for production, nonproduction, shared services, and inspection rather than one flat routing domain.
  • Peer regional TGWs only for approved prefixes and summarize VPC CIDRs to keep propagation and route review manageable.
  • Steer inter-VPC flows through regional inspection VPCs while keeping high-volume same-VPC traffic local.
  • Validate attachment and appliance throughput against the 10 Gbps profile, then publish per-path flow logs and utilization alarms.

Why interviewers ask this: The interviewer is evaluating route-domain design, centralized ownership, inspection, and throughput awareness at multi-account scale.

gatewaynetworking

I would choose Cloud WAN when global policy and segment lifecycle matter more than direct per-region route-table control.

  • Define segments for production, nonproduction, shared services, and regulated workloads in one core network policy.
  • Attach regional TGWs or VPCs through controlled automation and require explicit segment actions for cross-segment connectivity.
  • Use policy change sets to preview routing effects across all 5 regions before execution.
  • If the estate stays at 2 regions with one network team, TGW would be simpler and avoid Cloud WAN policy and data-processing overhead.

Why interviewers ask this: The interviewer is checking whether the service choice follows organizational scale and operating model rather than novelty.

hybrid-networkingdependencies

I would terminate the Direct Connect links in separate colocation facilities and on separate customer routers.

  • Use a Direct Connect Gateway with Transit Gateway associations so approved VPC routes are managed centrally across accounts.
  • Build site-to-site VPN tunnels over diverse internet providers and advertise them with less-preferred BGP paths.
  • Control route preference through BGP communities and explicit prefixes, avoiding accidental asymmetric paths through stateful firewalls.
  • Test each link and facility withdrawal quarterly while tracking failover convergence against a 60-second network objective.

Why interviewers ask this: The interviewer is evaluating physical diversity, BGP behavior, centralized attachment, and a measurable fallback design.

designdns

I would place Route 53 Resolver inbound and outbound endpoints in shared-services VPCs in at least 2 Availability Zones.

  • Forward only specific on-premises suffixes from AWS and conditionally forward AWS private namespaces from on-premises to inbound endpoints.
  • Use Route 53 Profiles or automated hosted-zone associations to distribute the 120 private zones to authorized VPCs.
  • Centralize Resolver query logs and detect NXDOMAIN spikes, forwarding loops, and endpoint capacity pressure.
  • Keep TTLs and deployment automation aligned to the 5-minute objective, while avoiding low TTLs on records that rarely change.

Why interviewers ask this: The interviewer is checking split-horizon DNS, conditional forwarding, multi-account association, and propagation trade-offs.

networkingprivate-connectivity

I would publish each service behind a Network Load Balancer as an endpoint service from dedicated provider accounts.

  • An Organizations ARN is not a valid endpoint-service allowed principal, so automation would expand the 40-account inventory into explicit account-root or approved role ARNs and reconcile additions and removals.
  • If broad discovery requires `*`, I would keep acceptance required and authorize each request against the account, role, service, and environment workflow before accepting it.
  • Consumers create interface endpoints in their own subnets, while verified private DNS and a central catalog map all 15 services to endpoint, port, owner, and SLA.
  • Monitor rejected and pending requests, NLB health, endpoint bytes, and hourly endpoint cost because 15 services across 40 accounts can create hundreds of endpoints.

Why interviewers ask this: The interviewer is evaluating PrivateLink address isolation and the exact principal and acceptance controls needed to authorize 40 consumer accounts.

designgatewaynetworking

I would deploy an egress and inspection VPC per region so a regional dependency does not pull internet traffic across regions.

  • Route workload default traffic through TGW appliance mode to redundant Gateway Load Balancer firewall endpoints.
  • Place NAT gateways after inspection in each Availability Zone and preserve zonal affinity to avoid cross-AZ data charges and asymmetric state.
  • Use SCPs and Config rules to block workload internet gateways, public IP assignment, and unapproved alternate egress paths.
  • Scale and load-test firewall fleets beyond 20 Gbps, then advertise only approved inter-region prefixes over TGW peering or Cloud WAN.

Why interviewers ask this: The interviewer is checking centralized inspection, symmetric routing, regional independence, and enforcement against bypass.

kubernetesawscompute

I would place each service by execution constraint rather than force all 12 onto one platform.

  • Use Lambda for short stateless burst handlers where duration, package, and concurrency limits fit the 20-fold spike.
  • Use ECS Fargate for containerized services that need long-running processes but not Kubernetes APIs or host control.
  • Choose EKS only if the organization already operates Kubernetes capabilities needed by several teams, and EC2 for kernel, GPU, or licensing constraints.
  • Compare p99 startup latency, steady utilization, operational staffing, and monthly cost in a load test before standardizing the default.

Why interviewers ask this: The interviewer is evaluating workload placement from concrete constraints rather than service preference.

I would choose EC2 because the kernel, memory, local NVMe, and licensing constraints require host-level control.

  • Select a memory-optimized or storage-optimized family only after benchmarking the engine against actual query and spill patterns.
  • Use dedicated hosts if the license is socket or host bound, and verify license mobility before committing spend.
  • Place instances in an Auto Scaling group for replacement while persisting durable state outside instance-store NVMe.
  • Use Capacity Reservations for required Availability Zones and Savings Plans only for the stable baseline that can survive family commitments.

Why interviewers ask this: The interviewer is checking whether hard host and licensing constraints override a default preference for managed containers.

kubernetesdesign

I would use separate AWS accounts and EKS clusters for regulated workloads because namespaces are not a hard tenant boundary.

  • For shared clusters, give each team a namespace with RBAC groups, ResourceQuota, LimitRange, and default-deny NetworkPolicies.
  • Use EKS Pod Identity or IRSA with one narrowly scoped IAM role per workload, never the worker-node role for application access.
  • Enforce allowed images, security contexts, and node placement through admission policy, with dedicated node pools for sensitive shared workloads.
  • Allocate control-plane, node, and observability costs by team labels while capping noisy-neighbor CPU, memory, and pod consumption.

Why interviewers ask this: The interviewer is evaluating the distinction between soft Kubernetes tenancy and a hard account and cluster boundary.

configkuberneteskubernetes-autoscaling

I would not promise both constraints uniformly: services with a hard 10% loss limit need at least 90% non-interruptible replicas, while flexible workloads carry extra Spot to reach 70% platform-wide.

  • Define Karpenter NodePools with diverse instance families, sizes, and Availability Zones rather than a narrow Spot fleet.
  • Spread critical replicas across zones and nodes, reserving on-demand capacity for at least 90% of each constrained service.
  • Use PodDisruptionBudgets for voluntary consolidation, but do not present them as protection from involuntary EC2 interruption.
  • Enable interruption handling and track pending time, replica loss, and Spot share before increasing flexible workloads toward the 70% target.

Why interviewers ask this: The interviewer is checking whether the candidate combines Karpenter economics with application-level disruption controls.

kuberneteslatencyservice-mesh

I would adopt a mesh only for capabilities that 120 services cannot reliably implement through simpler shared libraries or gateways.

  • Benchmark sidecar or ambient data-plane latency and CPU under production traffic, rejecting the design if p99 exceeds 10 ms.
  • Start with mTLS identity and traffic telemetry, not retries on every hop because layered retries amplify load.
  • Roll out to 10 low-risk services first and define escape hatches for protocols or workloads the mesh handles poorly.
  • Include certificate rotation, proxy upgrades, policy debugging, and on-call load in the capacity plan for the 4-person platform team.

Why interviewers ask this: The interviewer is evaluating whether a mesh decision includes latency, staffing, retry behavior, and an incremental adoption test.

deployment

I would provide versioned templates for an approved runtime, IaC, pipeline, telemetry, security checks, and ownership metadata.

  • A portal request creates the repository, account and environment bindings, least-privilege role, dashboards, alerts, and deployment workflow.
  • Expose paved-road extension points for data stores and queues instead of letting teams fork the entire template.
  • Publish measurable contracts such as 15-minute bootstrap, pipeline duration, supported versions, and platform support boundaries.
  • Track adoption, template divergence, failed deployments, and time to first production release across all 18 teams.

Why interviewers ask this: The interviewer is checking whether a golden path is a maintained product with contracts and escape hatches, not a repository skeleton.

deploymentrollback

I would use a progressive canary because the deployment frequency and 5-minute rollback target require automated evidence, not manual approval.

  • Send 2 percent of traffic to the new version, then advance through fixed stages only when error rate and p99 latency stay within guardrails.
  • Compare canary and baseline with CloudWatch or OpenTelemetry metrics and automatically reverse traffic on breached thresholds.
  • Keep database changes backward compatible using expand and contract so either application version can run during the canary.
  • Reserve blue-green for rare platform or dependency changes where duplicating the full environment is worth the faster traffic switch.

Why interviewers ask this: The interviewer is evaluating deployment strategy from SLA, release frequency, rollback time, and data compatibility.

databasedesign

I would run the primary Aurora cluster in the write region and a provisioned secondary cluster in the recovery region.

  • Use Global Database storage replication and monitor cross-region replication lag against the sub-second RPO.
  • Route reads to regional readers, but send all writes to the primary endpoint to preserve one authoritative writer.
  • Automate managed failover or switchover with Route 53 or application endpoint updates and rehearse completion within 60 seconds.
  • Keep schema, parameter groups, secrets, capacity, and application dependencies pre-provisioned identically in both regions.

Why interviewers ask this: The interviewer is checking whether the candidate connects Aurora topology and endpoint behavior to quantified RPO and RTO.

Locked questions

  • 21

    Aurora must serve users in 3 regions with local reads under 50 ms, but cross-region write conflicts are forbidden; what topology do you choose?

  • 22

    A Lambda API can reach 10,000 concurrent invocations, while Aurora permits only 800 application connections; how would you use RDS Proxy?

    proxyserverlessconcurrency
  • 23

    How would you design a DynamoDB global table for 200,000 writes per second across 3 regions when every item must have deterministic conflict handling?

    dynamodbdesign
  • 24

    One tenant produces 40% of 1,000,000 DynamoDB writes per second and tenant ID is the current partition key; how would you remove the hot partition?

    dynamodbpartitioning
  • 25

    DynamoDB Streams must feed 50,000 events per second to 8 consumers, but duplicate side effects are forbidden; what architecture do you propose?

    dynamodbarchitecture
  • 26

    How would you design S3 storage for 5 PB of regulated objects with 7-year retention, cross-region RPO under 15 minutes, and no administrator able to delete retained data?

    designobject-storageaws
  • 27

    An API serves 100,000 reads per second, allows data to be stale for 5 seconds, and updates through events; how would you combine ElastiCache and event-driven consistency?

    eventsconsistencyapi
  • 28

    How would you design zero-trust IAM for 800 people and 2,000 workloads across 60 accounts when human sessions are capped at 1 hour and long-lived access keys are forbidden?

    sessionszero-trustiam
  • 29

    How would you design KMS encryption for 60 accounts and 3 regions when regulated teams require key separation and central security must retain policy control?

    encryptiondesign
  • 30

    Ten thousand application secrets across 60 accounts must rotate every 30 days without exposing plaintext to the platform team; how would you use Secrets Manager?

    secrets
  • 31

    How would you aggregate GuardDuty, Security Hub, and Config for central security across all 60 accounts and enabled regions when new high-severity findings must be visible within 5 minutes?

    aggregationconfigseverity-priority
  • 32

    How would you design immutable audit logging for 60 accounts with 7-year retention, legal hold, and a hard constraint that workload administrators cannot alter logs?

    retentiondesignlogging
  • 33

    How would you design AWS supply-chain controls for a platform that builds 200 container images per day when production may run only signed images with traceable provenance?

    designcontainers
  • 34

    Twelve accounts process EU regulated data that must never leave eu-west-1 and eu-central-1, while shared tooling spans 3 global regions; how do you enforce the boundary?

    concurrency
  • 35

    For 60 accounts, 12 teams, and a hard requirement for reviewed plans before every change, how would you choose among Terraform Enterprise, AWS CDK, and CloudFormation?

    terraformiac
  • 36

    How would you design Terraform Enterprise state tenancy for 60 accounts and 400 workspaces when one team must not be able to plan against another team’s state?

    terraformdesign
  • 37

    Twenty teams need reusable IaC modules, and 80% of new services must follow the golden path within 6 months; how would you design the module program?

    designiac
  • 38

    Every IaC pull request for 60 accounts must receive policy-as-code results within 10 minutes; what controls do you put in the pipeline?

    ci-cdiaccode-review
  • 39

    How would you design a control that detects unauthorized IaC drift across 25,000 AWS resources within 15 minutes without automatically overwriting emergency changes?

    iacdesign
  • 40

    A new network baseline must roll out to 60 accounts in 5 waves, with a hard stop if more than 1 account fails validation; how do you make the rollout safe?

    validation
  • 41

    A customer API targets 99.95% availability over 30 days; how would you define its SLO and error-budget policy?

    sloapi
  • 42

    How would you design CloudWatch and OpenTelemetry telemetry for 200 services producing 1,000,000 spans per minute when observability spend is capped at 5% of AWS cost?

    observabilitymonitoringdesign
  • 43

    Which of backup-and-restore, pilot light, warm standby, and active-active would you choose for a tier-1 application requiring RTO 15 minutes and RPO 5 minutes across 2 regions?

    backups
  • 44

    How would you design backup immutability for 60 accounts with daily recovery points, 7-year regulated retention, and a hard constraint that compromised account admins cannot delete backups?

    designbackupsimmutability
  • 45

    How would you plan a quarterly game day for a 99.99% service with RTO 30 minutes, RPO 1 minute, and a hard constraint of no customer-visible interruption?

  • 46

    You must complete Well-Architected reviews for 40 workloads in 90 days without producing checklist-only reports; how would you run the program?

  • 47

    Monthly AWS spend is $1.2 million across 60 accounts, and 70% of compute usage is stable; how would you size Savings Plans and Reserved Instance commitments?

  • 48

    How would you allocate at least 95% of $1.2 million monthly AWS spend across 60 accounts and 18 product teams when shared network and platform costs are 20%?

  • 49

    Leadership mandates a 25% reduction from $900,000 monthly AWS spend within 6 months without lowering the 99.95% customer SLA; what is your concrete plan?

  • 50

    Twelve teams disagree on 3 AWS platform options, an RFC decision is due in 2 weeks, and 4 engineers you mentor must help set the technical direction; how do you lead the decision?

    conflictmentoringdecision-making
  • 51

    Route 53 sends 60% of 18,000 requests per second to us-east-1; CloudWatch is green, but 3 external probes show 38% checkout failures for 7 minutes after a health-check change. What routing decision do you make, and what evidence gates recovery?

    dnsmonitoring
  • 52

    A Transit Gateway change in a 220-account organization propagated 10.42.0.0/16 from development into production; VPC Flow Logs show 14,000 rejected flows and 6 payment services have been unreachable for 9 minutes. How do you contain the leak and decide routing is safe?

    gatewaynetworking
  • 53

    A 10 Gbps Direct Connect link dropped 12 minutes ago and BGP moved traffic to two 2 Gbps VPN tunnels, but packet captures show return traffic still using Direct Connect and 31% of API calls time out at 3.5 Gbps. What do you change, and what gate prevents another asymmetric path?

    hybrid-networkingapi
  • 54

    A checkout VPC has 48,000 concurrent outbound connections; NAT ErrorPortAllocation reaches 9,200 per minute, partner TLS timeouts hit 27%, and Flow Logs show one scraper opening 40 connections per second per task. How do you contain and gate egress recovery?

    resiliencenetworkingtls
  • 55

    An ALB surges from 24,000 to 61,000 requests per second after a promotion; target 5xx reaches 18%, logs show target_reset, and 640 ECS tasks use only 70% CPU. Do you scale, roll back, or change ALB, and how do you recover?

    rollback
  • 56

    After a Route 53 Resolver rule update, 73 VPCs cannot resolve corp.internal for 11 minutes; query logs show 82% SERVFAIL, the outbound endpoint ENIs are healthy, and on-premises DNS answers direct queries in 12 ms. What do you roll back, and what proves recovery?

    queriesdnsendpoints
  • 57

    A new PrivateLink endpoint service was created around a replacement NLB; 46 new endpoints await acceptance, while 160 existing endpoints still use the old service and see 34% resets after its NLB was changed. What do you restore, and how do you gate migration?

    migrationsendpointsasync
  • 58

    A routing change sends 94 TB per day of EU-regulated customer traffic through us-east-1 for 27 minutes, violating the approved residency boundary; live flow telemetry shows 38 Gbps crossing regions. How do you contain the incident, preserve evidence, and recover compliantly?

    incidents
  • 59

    A mixed EC2 fleet loses 78% of Spot capacity in 6 minutes during a 32,000 requests-per-second peak; 420 instances terminate, queue depth reaches 2.8 million, and notices give 2 minutes. How do you recover without another overload?

    capacitycomputedata-structures
  • 60

    An Auto Scaling rollout replaced 180 of 240 instances with a launch template whose user data fails; EC2 console shows cloud-init exit 1, ALB healthy hosts fell to 54, and errors reached 23% in 8 minutes. What do you roll back and how do you replace safely?

    scalingcomputerollback
  • 61

    An ECS deployment of 320 Fargate tasks reaches steady state, yet checkout 5xx climbs from 0.3% to 11%; memory is 58%, the circuit breaker did not fire, and traces identify the new image digest. What do you do and what recovery gate was missing?

    deploymentcontainersresilience
  • 62

    An EKS cluster with 900 nodes has 19% failed requests; Kubernetes API p99 is 14 seconds, 37 nodes are NotReady, and AWS Health is clear after a CNI rollout 6 minutes earlier. Do you replace nodes, roll back, or fail over, and what proves recovery?

    kubernetesrollbackapi
  • 63

    Karpenter consolidation drains 140 EKS nodes in 9 minutes despite a 10% disruption budget; 2,300 pods are Pending, payment p99 reaches 4.2 seconds, and logs show incompatible topology constraints. What do you stop, and how do you restore scheduling?

    kuberneteskubernetes-autoscalingjobs
  • 64

    After an App Mesh sidecar update on 85 services, p99 rises from 180 ms to 1.9 seconds and retries triple traffic to 72,000 requests per second; X-Ray shows 640 ms in Envoy and application CPU is 46%. Do you tune or roll back, and how do you recover?

    rollback
  • 65

    A retry bug drives Lambda from 8,000 to 54,000 invocations per second; concurrency hits 95,000, throttles affect 14 unrelated functions, and SQS oldest age reaches 41 minutes. What concurrency decision do you make and how do you drain safely?

    resilienceserverlessqueues
  • 66

    ECR scanning flags a critical RCE CVE in a base image used by 1,700 running containers across 26 services; exploitation evidence is absent, but the patched image fails 9% of canary health checks. Do you stop production, accept risk, or roll back, and what gates recovery?

    containershealth-checksrollback
  • 67

    An Aurora PostgreSQL writer fails over in 38 seconds, but 1,200 tasks reconnect at once; connections reach 14,800 of 15,000, CPU is 96%, and checkout errors remain 17% after 6 minutes. How do you stop the storm and decide the writer is safe?

    postgres
  • 68

    RDS PostgreSQL has 90 GB free on a 12 TB gp3 volume, space falls 18 GB per hour, replica lag is 27 minutes, and Performance Insights ties growth to a failed index build started 4 hours ago. Do you add storage, cancel, or fail over, and how do you recover?

    indexespostgresreplication
  • 69

    An RDS MySQL replica serving 68% of catalog reads falls 46 minutes behind while primary writes reach 21,000 per second; DiskQueueDepth is 190 and stale prices cause 7% cart mismatches. Do you promote, stop reads, or rebuild, and what gates recovery?

    replicationrds
  • 70

    A DynamoDB orders table receives 180,000 writes per second, but merchant-8842 generates 42%, roughly 75,600 writes per second; throttles reach 31,000 per second and the merchant appears in 87% of failures. How do you contain the hot key without losing order sequence?

    dynamodbthrottle
  • 71

    A DynamoDB global table accepts writes in us-east-1 and eu-west-1 during a 13-minute partition; replication resumes, but 26,400 profiles have conflicting addresses and CloudTrail confirms both regions stayed writable. Which copy wins, and how do you recover?

    replicationdynamodbpartitioning
  • 72

    An operator deletes 18 million objects from a versioned S3 bucket; 6.2 million delete markers already replicated, Object Lock is absent, and downloads fail at 44% after 17 minutes. What do you restore first, and how do you repair the destination?

    replicationawsobject-storage
  • 73

    An ElastiCache Redis primary fails over in 54 seconds, but 700 instances miss reconnect and hit rate falls from 93% to 8%; Aurora CPU reaches 94% and checkout p99 is 3.8 seconds. Do you flush, fail back, or throttle, and how do you prevent a stampede?

    rediscachingthrottle
  • 74

    An SQS order queue grows from 40,000 to 12.6 million messages in 35 minutes; oldest age is 29 minutes, Lambda errors are 0.2%, and Aurora write latency rises from 12 ms to 140 ms above 4,000 concurrency. How do you drain and define recovery?

    concurrencydata-structureslambda
  • 75

    GuardDuty reports an IAM key used from 3 countries in 11 minutes; CloudTrail shows 286 AssumeRole calls and 41 security-group changes, while the owner says the key is idle. How do you contain compromise and decide accounts are safe?

    iam
  • 76

    A new SCP has a wrong condition across 310 accounts, and 74 production roles lose AWS API access; CloudTrail records attachment at 14:06 and break-glass roles also fail. How do you recover from lockout, and what gate allows policy reapplication?

    linuxapi
  • 77

    A KMS customer key used by 96 services is disabled by automation; 63% of decrypt calls fail, CloudTrail shows DisableKey 8 minutes ago, and 4 TB of queued records use that key. Do you reenable, rotate, or restore, and how do you prove recovery?

    encryptiondata-structures
  • 78

    Secrets Manager rotates a database password used by 380 tasks, but clients cache it for 24 hours; authentication failures reach 71%, the old secret is valid for 5 more minutes, and finishSecret succeeded. What do you roll back, and how do you recover?

    authsecretspasswords
  • 79

    GuardDuty raises S3 exfiltration for 7.4 TB read by an EC2 role in 2 hours; Flow Logs show 3.1 TB sent to one unfamiliar IP, while normal reads are 20 GB per day. Do you terminate, block, or preserve, and what gates recovery?

    computeobject-storageaws
  • 80

    Security finds a 19-hour CloudTrail gap across 84 accounts; S3 delivery stops at 02:11, CloudTrail Lake has events for 61 accounts, and an IAM incident may have occurred at 08:40. How do you preserve evidence and decide audit coverage is restored?

    incidentscoverageaws
  • 81

    AWS Config reports a public S3 bucket with 2.3 million payroll objects; logs show 18,400 anonymous GETs over 36 minutes, Block Public Access was disabled at 10:02, and legal needs a decision in 15 minutes. What do you do and when can service resume?

    configawsobject-storage
  • 82

    EKS audit logs show a service account creating cluster-admin bindings in 4 clusters; 63 privileged pods launch in 22 minutes, 9 nodes contact a mining pool, and the token should only list pods. How do you contain escalation and rebuild trust?

    kubernetestokensescalation
  • 83

    A Terraform Enterprise workspace loses its state lock during apply; 320 of 500 resources changed, the latest state is 14 MB smaller, and the next plan deletes 86 live resources. What do you restore, and what evidence permits another apply?

    terraform
  • 84

    A shared Terraform module update across 74 workspaces changes a default from private to public and plans replacement of 1,900 load balancers; 11 already applied and Config reports 7 public endpoints. How do you stop blast radius and recover consumers?

    terraformconfigload-balancing
  • 85

    A CloudFormation update to a 260-resource stack fails at resource 214 with UPDATE_ROLLBACK_FAILED; 37 resources changed, database replacement is blocked, and API errors are 28%. Do you continue rollback, import, or restore, and what gates recovery?

    rollbackiacapi
  • 86

    A CDK deployment succeeds for 48 stacks, but drift detection finds 312 IAM and security-group differences after an emergency change; 6 roles have AdministratorAccess and no outage is active. Do you overwrite drift or preserve it, and what gate do you use?

    deploymentiaciam
  • 87

    A pipeline role creates 23 IAM users and disables 16 CloudWatch alarms in 9 minutes; its session came from one build, and 140 production accounts trust it. How do you contain the CI role compromise and restore delivery?

    sessionsiamci-cd
  • 88

    During a 21-minute checkout outage, dashboards stay green because metric ingestion stopped in one account; ALB logs show 36% 5xx, collectors drop 84 million spans, and the status page claims 99.99%. How do you operate blind and decide monitoring is trustworthy?

    monitoring
  • 89

    A new alarm package creates 48,000 notifications in 3 hours across 260 accounts; only 6 map to customer impact, acknowledgment p95 reaches 19 minutes, and one real database alarm is missed. What do you silence, and what evidence gates reactivation?

    reactdatabase
  • 90

    A canary sends 5% traffic to a release for 12 minutes; p99 rises from 260 ms to 1.4 seconds, errors remain 0.4%, and the gate passes because it checks only 5xx below 1%. Do you promote or roll back, and how do you repair the gate?

    deployment-strategiesrollback
  • 91

    us-east-1 is unavailable for 7 minutes for payments with RTO 10 minutes and RPO 30 seconds; Aurora Global Database lag is 18 seconds, standby capacity is 40%, and Route 53 still sends 22% east. Do you activate DR now, and what gates recovery?

    capacitydnsdatabase
  • 92

    A restore test must recover a 9 TB Aurora snapshot within a 4-hour RTO, but after 3 hours only 22% of hot data is readable and 31% of validation queries time out; AWS Backup reports completion. Do you accept, switch methods, or declare failure, and how do you recover?

    queriessnapshotvalidation
  • 93

    AWS daily run rate rises 40% from $82,000 to $114,800 in 6 hours; Cost Explorer attributes $21,000 to NAT bytes and $9,600 to us-east-1 transfer after a release, while customer traffic is flat. What do you stop, and what evidence permits normal operation?

    networking
  • 94

    A $1.8 million annual Compute Savings Plan covers 72% of baseline, but demand has fallen 35% for 3 months and utilization is 61%; finance asks today whether to buy another 1-year commitment for a launch. What is your decision and what evidence could reverse it?

  • 95

    A regional capacity event removes 600 EC2 instances; the recovery Auto Scaling group requests 600 c7i.4xlarge instances, but only 240 of the regional Standard On-Demand vCPU quota remain and launches fail with VcpuLimitExceeded. How do you contain the quota incident and gate recovery?

    scalingcapacitycompute
  • 96

    Control Tower detects 186 drifted resources across 42 of 500 accounts after a landing-zone update; 17 mandatory log buckets fell from required SSE-KMS to default SSE-S3 and received 2.8 million objects, 9 controls are disabled, and vending continues at 12 accounts per hour. What do you stop and how do you recover governance?

    iacawslanding-zone
  • 97

    An external PCI audit is due in 72 hours, but Security Hub shows 1,840 failed controls across 230 accounts; 93% come from one new standard, 14 expose card-data paths, and automation fails 6%. What do you fix first, and what evidence supports compliance?

  • 98

    A production incident affects 38% of checkouts for 24 minutes; 46 people join the bridge, 7 teams make simultaneous changes, and no one owns customer updates while errors rise from 12% to 19%. How do you take command and what recovery gate do you enforce?

    incidentsjoins
  • 99

    A mid-level engineer proposes opening all security-group ingress to 0.0.0.0/0 for 15 minutes to restore a service with 62% timeouts; Reachability Analyzer shows one missing port 8443 rule, 12 teams share the VPC, and RTO has 9 minutes left. How do you mentor the decision and recover safely?

    resiliencenetworkingmentoring
  • 100

    A 67-minute outage caused $2.4 million in failed orders; the trigger was an ALB timeout change, but evidence also shows a 19-minute alert delay, no rollback owner, and a DR test skipped for 8 months. You can fund only 2 actions this quarter. What do you prioritize and how do you verify recovery capability?

    rollbackalertingprioritization