Skip to content

Site Reliability Engineer interview questions

100 real questions with model answers and explanations for Senior candidates.

See a Site Reliability Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

designsloavailability

I would run the complete request path across three zones and size any two zones to carry peak traffic.

  • A 99.99% monthly target allows about 4.32 minutes of bad availability, so zone recovery cannot depend on manual provisioning.
  • Stateless services, ingress, caches, queues, and the database quorum all span zones without one shared egress or identity dependency.
  • Each zone provides roughly half of peak capacity, keeping normal peak utilization near 67% before one zone disappears.

Why interviewers ask this: The interviewer checks whether the candidate translates an availability target into failure domains, dependency redundancy, and capacity math.

disaster-recovery

I would choose active-passive write ownership because global ordering and RPO zero require one fenced serialization path.

  • The passive region continuously receives synchronous or quorum-protected data, accepting WAN latency or reduced write availability during partition.
  • Reads may run active-active if their freshness contract allows it, but only one region holds the current fencing token for writes.
  • A timed promotion test must fence the old writer and restore writes inside ten minutes before the design meets its RTO claim.

Why interviewers ask this: A strong answer chooses a topology from ordering, data-loss, and recovery requirements rather than preferring active-active by default.

failoverapi

I would require sustained user-path failure from independent observers and proof that the destination is ready.

  • At least three of five probes from separate networks must see over 5% errors or SLO-breaking latency for two minutes.
  • The target region must have healthy dependencies, replication within RPO, and capacity for the full 40,000 requests per second.
  • Failover uses hysteresis and a 15-minute minimum hold period so brief recovery cannot flap traffic between regions.

Why interviewers ask this: The interviewer evaluates failure detection, false-positive control, and destination readiness with explicit thresholds.

replicationlatencydisaster-recovery

I would use asynchronous cross-region replication because synchronous acknowledgement cannot fit a 25 ms budget over an 80 ms link.

  • Local writes acknowledge after durable regional commit, while replication lag is measured against the 30-second RPO.
  • Alerts fire before lag consumes the full budget, and promotion is blocked when the replica position is outside it.
  • Ledgers that truly require zero loss need a different latency contract or a quorum placement closer to writers.

Why interviewers ask this: A strong answer makes the latency-versus-data-loss decision explicit and avoids claiming incompatible guarantees.

databasedistributed-systems

I would require a consensus-backed lease and fence the former primary before the new region accepts writes.

  • Promotion obtains a monotonic fencing token from an odd quorum placed across three independent failure domains.
  • Storage and downstream consumers reject writes carrying an older token, even if the old primary is still reachable by some clients.
  • Loss of quorum stops promotion and write availability; it never silently weakens the single-writer invariant to meet five minutes.

Why interviewers ask this: The interviewer checks consensus, fencing, and willingness to sacrifice write availability rather than corrupt ordered state.

designdisaster-recoverykubernetes

I would run autonomous clusters per region and pre-provision the recovery region for the declared traffic shift.

  • Each cluster has local ingress, DNS, registry, secrets, observability, and no control-plane dependency on the other region.
  • The recovery region keeps enough warm capacity to reach 60,000 requests per second inside 15 minutes after cache warm-up.
  • Global traffic moves only after data fencing and dependency checks, while GitOps keeps reviewed configuration equivalent without one shared runtime controller.

Why interviewers ask this: A strong answer preserves regional autonomy and includes capacity, data, and platform dependencies in the RTO.

design

I would create at least 20 repeatable cells and route each user deterministically to one cell.

  • Every cell owns bounded compute, queues, caches, and data partitions so overload cannot consume a healthy cell's pool.
  • Global identity and routing stay minimal and redundant because a shared dependency could erase the 100,000-user blast-radius goal.
  • Capacity planning includes one spare cell and a controlled tenant-rebalancing path rather than moving users during an active failure.

Why interviewers ask this: The interviewer evaluates blast-radius arithmetic, dependency isolation, and the operating cost of cell recovery.

secrets

I would use a regional Vault performance replica with tested promotion and leases that outlive the 30-minute isolation window.

  • Applications authenticate through regional workload identity and cache only current credentials, not reusable root or bootstrap tokens.
  • Promotion includes seal or auto-unseal availability, storage health, token semantics, and fencing of the former active cluster.
  • A game day proves all 80 services renew or degrade safely for 30 minutes without bypassing authorization.

Why interviewers ask this: A strong answer covers Vault replication, credential lifetime, promotion, and application behavior during isolation.

I would first challenge total global order and shard serialization by account if the invariant permits it.

  • Per-account leaders preserve the order users need while independent shards process unrelated accounts in parallel.
  • If total order is mandatory, one consensus leader and WAN quorum become the throughput and latency ceiling.
  • A benchmark must sustain 50,000 writes per second through leader loss and show the exact write unavailability during quorum changes.

Why interviewers ask this: The interviewer checks whether the candidate narrows an expensive ordering guarantee and validates the unavoidable coordination cost.

designfailoverdisaster-recovery

I would run a full traffic and data exercise from external probes, not only stop application instances.

  • The test removes regional ingress and the replication link, then measures detection, decision, promotion, traffic shift, and cache warm-up separately.
  • Database positions and business checks prove data loss stays within 60 seconds and old-primary fencing prevents dual writes.
  • The design passes only after two consecutive exercises restore user success within 12 minutes and failback is also rehearsed.

Why interviewers ask this: A strong answer validates end-to-end recovery and data correctness with repeated timed evidence.

capacitycapacity-planning

The baseline forecast is about 57,800 requests per second, then I add failure and forecast headroom.

  • The calculation is 25,000 multiplied by 1.15 to the sixth power, not a flat 90% addition.
  • I benchmark sustainable throughput at the latency SLO and plan at least 30% uncertainty above the forecast.
  • Zone-loss capacity is calculated separately, so normal placement cannot consume the headroom reserved for the largest tolerated failure.

Why interviewers ask this: The interviewer evaluates growth compounding, performance limits, uncertainty, and failure capacity as separate inputs.

scalingcapacitycapacity-planning

I would provision each zone for 15,000 requests per second, giving 45,000 total and 30,000 after one zone fails.

  • Normal peak uses about 67% of total capacity, which is the cost of immediate one-zone tolerance.
  • The calculation includes database connections, queue throughput, egress, and cache capacity, not only application CPU.
  • A zone-loss load test must hold latency and errors at 30,000 requests per second before the headroom is considered real.

Why interviewers ask this: A strong answer performs the failure-capacity math and applies it to every constrained dependency.

scalingautoscaling

I would keep warm capacity for the 60-second burst because node autoscaling cannot beat a four-minute startup.

  • Workload scaling uses a leading signal such as queue age or concurrency before CPU saturation.
  • Placeholder capacity or a minimum node floor covers the measured burst while new nodes join.
  • I separately measure cloud launch, bootstrap, CNI, image pull, and warm-up to reduce four minutes without pretending it is immediate.

Why interviewers ask this: The interviewer checks whether the candidate recognizes a timing mismatch and combines predictive scaling with reserved headroom.

capacitycapacity-planning

I would size requests per workload from sustained percentiles and burst behavior, not the cluster-wide 35% average.

  • CPU requests near p95 plus workload-specific margin improve placement, while latency-sensitive bursts may use uncapped spare CPU.
  • Memory uses a higher working-set percentile because underestimation causes eviction or OOM rather than throttling.
  • I canary new values and retain enough allocatable capacity for one-zone loss while checking throttling, Pending Pods, and SLOs.

Why interviewers ask this: A strong answer distinguishes averages, tail usage, CPU and memory behavior, and failure headroom.

concurrencydata-structures

I need at least 13 workers during recovery because 400 per second drains the burst while 100 per second handles new arrivals.

  • The drain requirement is 240,000 divided by 600 seconds, then incoming traffic is added before dividing by worker throughput.
  • I would provision about 16 workers for retries and processing variance, subject to broker and downstream limits.
  • Queue age, not only depth, drives scaling and confirms the oldest message falls below ten minutes.

Why interviewers ask this: The interviewer checks queue-drain arithmetic and whether new arrivals, variability, and downstream capacity are included.

postgresshardingreplication

I would optimize and scale the primary vertically first because read replicas do not remove a write bottleneck.

  • Query profiles, batching, indexes, connection overhead, and WAL or storage limits show whether CPU is the true constraint.
  • A larger primary buys time with the least application complexity, but I document its hardware and failover ceiling.
  • Sharding starts only when projected writes exceed that ceiling, using a key that spreads load and preserves required transaction boundaries.

Why interviewers ask this: A strong answer matches the scaling method to a measured write bottleneck and treats sharding as a deliberate complexity cost.

designapiload-testing

I would reproduce the production endpoint mix with an open workload model and test beyond both peak and failure capacity.

  • Payload sizes, cache state, geographic latency, authentication, and background jobs match production rather than uniform empty requests.
  • Load steps through 50%, 100%, and 130%, then repeats peak after removing one zone.
  • The test records latency histograms, errors, saturation, queues, and dependency limits long enough to expose leaks and autoscaling delay.

Why interviewers ask this: The interviewer evaluates representative workload design, coordinated-omission avoidance, and explicit failure-capacity validation.

capacitycapacity-planning

I would plan regional reservation ranges six weeks ahead and maintain a portable overflow tier for forecast error.

  • Demand is split into committed baseline, launch-driven upside, and batch work that can move in time or region.
  • Each region tracks quota, reserved GPUs, allocatable GPUs, queue age, and the lead time to add the next block.
  • I reserve the critical baseline plus measured uncertainty, while interruptible batch absorbs the 40% range instead of duplicating it everywhere.

Why interviewers ask this: A strong answer combines long provisioning lead time, uncertain demand, regional constraints, and movable workloads.

I would measure successful eligible checkouts and completion below a user-relevant latency threshold at the transaction boundary.

  • Availability counts committed orders as good and 5xx, invalid responses, or lost confirmations as bad.
  • Latency counts eligible journeys completed within 800 ms, while fast errors remain bad outcomes.
  • Dependency metrics explain failures, but retries and internal calls do not inflate the 10,000-request-per-second user denominator.

Why interviewers ask this: The interviewer checks precise good and valid events, latency semantics, and measurement at the user boundary.

batchslo

I would assign the strictest target to the class with the highest user and correctness cost, not the highest volume.

  • Interactive writes may receive 99.99% availability with a 500 ms latency threshold.
  • Interactive reads can use 99.9% if caching or stale fallback preserves value, while batch completion may use 99% within a deadline.
  • Every request maps to one class so high-volume reads cannot hide failure of low-volume critical writes.

Why interviewers ask this: A strong answer segments SLOs by business impact and prevents aggregate traffic from masking critical paths.

Locked questions

  • 21

    An API depends on three services each claiming 99.9%. Why is multiplying them into 99.7% not a sufficient service SLO design?

    designsloapi
  • 22

    A streaming endpoint legitimately runs for 30 seconds, while normal API calls should finish in 400 ms. How would the latency SLI handle both?

    endpointslatencystreaming
  • 23

    Calculate the monthly error budget for 99.9% availability at 100 million valid requests.

    reliabilityavailabilityerror-budget
  • 24

    Design multi-window burn-rate alerts for a 99.9% SLO over a 30-day window.

    designsloalerting
  • 25

    A critical admin API receives only 200 requests per day. How would you define a meaningful 99.9% reliability objective?

    reliabilityapi
  • 26

    Search returns cached results up to five minutes old during overload. Should those requests count as good in a 99.9% SLO?

    slocaching
  • 27

    A checkout journey has a 99.95% SLO and depends on payment, inventory, and identity. How would you set internal objectives?

    slo
  • 28

    Design OpenTelemetry collection for 500 services across three regions at 1 million telemetry events per second.

    designobservability
  • 29

    How would you operate Prometheus for 30 clusters and 15 million active series with 13 months of retention?

    retentionmonitoring
  • 30

    Metric volume grows from 4 million to 20 million series after teams add user_id and raw URL labels. What controls do you design?

    designmonitoring
  • 31

    200 services emit metrics, logs, and traces, but a release cannot be followed across all three. What correlation contract would you set?

    correlationmonitoring
  • 32

    A tracing platform receives 2 million spans per second but can retain only 5%. How would you sample?

  • 33

    Design observability to remain useful when traffic reaches 3x normal and telemetry volume threatens the collectors.

    designobservability
  • 34

    A platform has 800 services and currently sends 600 pages per day. What alerting design would you implement?

    alertingdesign
  • 35

    Grafana has 200 dashboards, and core incident views take 18 seconds to load. How would you redesign them?

    incidentsincident-managementmonitoring
  • 36

    Define severity levels for a platform serving 80 services and 5 million users.

    severity-priority
  • 37

    Design a sustainable 24/7 on-call rotation for ten SREs supporting 60 services.

    on-calldesign
  • 38

    A sev1 bridge attracts 30 engineers in ten minutes. What incident-command structure should already exist?

    incidentsincident-management
  • 39

    How would you design follow-the-sun incident handoff across three regions during an eight-hour sev1?

    incidentsincident-managementdesign
  • 40

    A postmortem program creates 120 actions per quarter but closes only 25%. How would you redesign the process?

    incidentspostmortemsconcurrency
  • 41

    A payment provider may be unavailable for 20 minutes, while your API must answer within two seconds. What degraded mode do you design?

    designapi
  • 42

    A multi-region API serves 100,000 requests per second, while one tenant is limited to 1,000. How would you divide enforcement across regions?

    api
  • 43

    A service is safe at 60,000 requests per second but may receive 90,000. How would you shed load?

  • 44

    A four-stage pipeline receives 8,000 events per second, but stage three can process only 6,000. How do you design backpressure?

    designresiliencebackpressure
  • 45

    A service makes 50,000 dependency calls per second. How would you cap retries at 10% extra load?

    dependencies
  • 46

    Set circuit-breaker behavior for a dependency that normally handles 20,000 calls per second at 0.2% errors.

    dependencies
  • 47

    How would you validate load shedding, retry budgets, and circuit breakers together at 80,000 requests per second?

    resiliencevalidationload-management
  • 48

    Design the first chaos experiment for a three-zone service with 99.95% availability and no prior game days.

    designavailabilitychaos-engineering
  • 49

    Plan a regional game day for RTO 15 minutes and RPO one minute without risking all customers.

    disaster-recovery
  • 50

    Forty tier-1 services need recurring chaos coverage. What program would you run over 12 months?

    chaos-engineeringcoverage
  • 51

    At 09:02 UTC, checkout 5xx jumps from 0.2% to 18% two minutes after release 4.17 reaches 40% of Pods; what do you do in the first 10 minutes?

  • 52

    At 14:10 UTC, a payment dependency slows from 80 ms to 3 seconds and three client layers amplify 8,000 requests per second into 54,000; how do you contain the retry storm?

    resiliencedependencies
  • 53

    Between 16:00 and 16:12, API p50 remains 45 ms but p99 rises from 320 ms to 4.8 seconds for 7% of users in one zone; how do you diagnose it?

    api
  • 54

    At 11:25 UTC, PostgreSQL reaches 500 of 500 connections, API errors hit 22%, and CPU is only 38%; what is your incident decision?

    incidentsincident-managementapi
  • 55

    At 08:40 UTC, Kubernetes cannot create nodes because the cloud account has used 2,000 of 2,000 vCPUs while 180 Pods remain Pending; how do you recover?

    kubernetes
  • 56

    At 03:15 UTC, an Elasticsearch data node is 97% full, watermarks block writes, and log ingestion is 140 GB per hour; what do you do before disk reaches 100%?

    search
  • 57

    At 19:00 UTC, Kafka consumer lag grows from 20,000 to 9 million records after one downstream shard slows to 400 events per second; incoming traffic is 6,000 per second. What do you do?

    shardingkafka
  • 58

    At 12:01 UTC, a DNS change lowers successful lookups from 99.99% to 91% for Android clients with a 30-minute TTL; what is your recovery plan?

    recoverydns
  • 59

    At 06:50 UTC, 28% of clients reject a renewed TLS certificate because one intermediate chain is missing, and the old certificate expires in 40 minutes; what do you do?

    tls
  • 60

    At 02:20 UTC, Redis fails over in 18 seconds, but 600 application Pods reconnect simultaneously and drive CPU to 96% with 35% timeouts; how do you stabilize it?

    redisresilience
  • 61

    A release at 10:00 burns 18% of a 30-day error budget in 25 minutes, but only a new reporting feature is failing; do you roll back the whole release?

    error-budgetreliabilityrollback
  • 62

    At 17:30 UTC, enabling a recommendation flag for 50% of users raises checkout p99 from 700 ms to 1.9 seconds without increasing errors; what do you decide?

  • 63

    At 13:05 UTC, a schema migration has held an ACCESS EXCLUSIVE lock for 95 seconds, 1,800 writes are queued, and replication lag is 70 seconds; what do you do?

    schemamigrationsreplication
  • 64

    At 09:45 UTC, a 2 TB cache expires at once and origin traffic rises from 15,000 to 110,000 requests per second; database CPU reaches 89%. How do you stop the stampede?

    zero-to-onedatabasecaching
  • 65

    At 18:20 UTC, a shipping provider starts returning 429 to 70% of calls, while 24,000 orders are waiting and its quota resets in 35 minutes; what do you do?

  • 66

    At 04:12 UTC, a Kubernetes node upgrade evicts 220 Pods, 65 remain Pending, and API availability falls to 96%; what is your response?

    availabilityapikubernetes
  • 67

    From 01:00 to 05:00, service memory grows 120 MB per hour per Pod; 14 of 80 Pods have OOM-killed and p99 is 1.4 seconds. How do you manage the leak overnight?

    memory
  • 68

    At 15:40 UTC, a Java service shows 9-second stop-the-world pauses every 2 minutes after heap occupancy reaches 82%, but CPU is 45%; what do you do?

    data-structures
  • 69

    At 20:05 UTC, one availability zone shows 4% packet loss, gRPC retries triple traffic to 72,000 calls per second, and the provider has no ETA; what do you do?

    availabilitygrpc
  • 70

    At 10:30 UTC, 12 of 100 API Pods receive 46% of traffic and hit 95% CPU while the rest stay below 30%; how do you handle the imbalance?

    restsoft-skills
  • 71

    At 21:15 UTC, one tenant sends 18,000 requests per second against a contracted 2,000, raising shared API p99 to 2.2 seconds for 300 tenants; what do you do?

    api
  • 72

    At 07:10 UTC, NTP drift reaches 95 seconds on 32 nodes, causing 14% of OAuth tokens to appear expired; how do you recover without accepting invalid tokens?

    oauthtokensiac
  • 73

    At 05:55 UTC, the identity provider rotates signing keys, 38% of API requests fail JWT validation, and the JWKS endpoint is timing out for 6 seconds; what do you do?

    jwtvalidationendpoints
  • 74

    At 09:20 UTC, a release adds customer_id to Prometheus labels, active series jump from 6 million to 48 million, and collectors start dropping 30% of samples; what do you do?

    monitoring
  • 75

    At 23:40 UTC, Grafana and tracing are unavailable during a 12% API error spike, but application logs and cloud metrics remain; how do you lead diagnosis for 20 minutes?

    monitoringapi
  • 76

    After a 25-minute outage, 4.5 million jobs are queued, normal capacity is 8,000 per second, and new arrivals are 6,000 per second; how do you recover without causing a second outage?

    capacitydata-structurescapacity-planning
  • 77

    At 18:00 UTC, traffic reaches 140,000 requests per second against a tested safe limit of 100,000, and checkout errors climb to 9%; which load do you shed?

  • 78

    At 10:05 UTC, mobile clients retry every 100 ms after 503 responses, turning 20,000 user actions per second into 160,000 requests; no client update can ship for 24 hours. What do you do?

    resilience
  • 79

    At 03:00 UTC, the primary region is down, the replica is 75 seconds behind, and 11,000 orders may be missing; the business asks for immediate failover. What do you decide?

    replicationfailover
  • 80

    At 13:40 UTC, reconciliation finds 620 duplicate charges created during a 17-minute timeout incident; service is now healthy. What do you do next?

    incidentsincident-managementreact
  • 81

    A 42-minute outage began at 09:07 when an engineer changed an Envoy timeout from 2 seconds to 200 ms; rollback took 31 minutes because ownership was unclear. How do you run the postmortem?

    ownershipincidentspostmortems
  • 82

    The same queue overflow caused 3 outages in 6 months despite a completed runbook action; the latest impact lasted 28 minutes. What do you change after review?

    runbooksdata-structures
  • 83

    Four SREs spend 26 hours per week manually restarting failed batch jobs; each automation attempt would take about 120 engineering hours. Do you automate now?

    batch
  • 84

    Two engineers spend 9 hours per week renewing 240 certificates, and one missed renewal caused a 16-minute outage; what automation do you approve?

  • 85

    The on-call SRE spends 14 hours per week granting temporary production access, with 60% of requests arriving after hours; what do you change?

    on-call
  • 86

    A release engineer spends 18 hours per week running 45 manual deployment steps, and 7% of releases need correction; what is your automation sequence?

    deployment
  • 87

    A team spends 11 hours per week investigating 320 alerts, but only 24 require action; what do you do during the next 2 weeks?

    alerting
  • 88

    During a 12-hour shift, the primary receives 17 pages, including 6 after midnight, and misses one 10-minute acknowledgement target; what do you do before the next shift?

  • 89

    At 02:10 UTC, one on-call engineer receives a database Sev1 and a separate CDN Sev1 within 3 minutes; only 4 responders are available. How do you allocate them?

    on-calldatabase
  • 90

    A Sev1 has run for 7 hours; the outgoing team has 25 minutes to hand over, with 3 mitigations active and replication lag at 48 seconds. How do you transfer command?

    replication
  • 91

    A junior on-call engineer sees errors rise from 0.3% to 6% after a 10% canary but wants 20 more minutes of diagnosis before rollback; how do you mentor in the moment?

    mentoringon-callrollback
  • 92

    A mid-level SRE manually raises a PostgreSQL connection limit from 400 to 800 during 2 incidents, worsening both; how do you coach and control the next incident?

    incidentsincident-managementpostgres
  • 93

    A new incident commander spends 15 minutes debugging one Pod while 12 responders receive no assignments during a 9% outage; how do you intervene and mentor?

    mentoringincidentsincident-management
  • 94

    Traffic grows 12% weekly for 5 weeks, database storage has 9 days left, and adding a shard takes 14 days; what action do you take today?

    databasesharding
  • 95

    During a traffic spike, finance asks you to remove 30% reserved compute to save $18,000 per day while one-zone headroom is only 12%; what do you decide?

  • 96

    A rollback from version 7.4 to 7.3 lowers errors from 14% to 4% but then stalls because 7.3 cannot read 8% of newly written records; what do you do?

    rollback
  • 97

    At 22:00 UTC, an L7 attack raises traffic from 25,000 to 900,000 requests per second, WAF CPU is 88%, and 3% of legitimate logins fail; how do you respond?

  • 98

    At 01:30 UTC, ransomware indicators appear on 6 build agents while a production hotfix is due in 20 minutes; security isolates CI. How do you keep the outage response safe?

  • 99

    At 11:11 UTC, an inventory call slows to 1.5 seconds, checkout has a 2-second deadline, and 4 sequential retries push p99 to 7 seconds; what do you change during the incident?

    incidentsestimationincident-management
  • 100

    After a 63-minute outage, errors have been below 0.2% for 8 minutes, but 1.2 million messages remain queued and replication lag is 35 seconds; when do you declare recovery?

    replicationdata-structures