Site Reliability Engineer interview questions
100 real questions with model answers and explanations for Senior candidates.
See a Site Reliability Engineer resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would run the complete request path across three zones and size any two zones to carry peak traffic.
- A 99.99% monthly target allows about 4.32 minutes of bad availability, so zone recovery cannot depend on manual provisioning.
- Stateless services, ingress, caches, queues, and the database quorum all span zones without one shared egress or identity dependency.
- Each zone provides roughly half of peak capacity, keeping normal peak utilization near 67% before one zone disappears.
Why interviewers ask this: The interviewer checks whether the candidate translates an availability target into failure domains, dependency redundancy, and capacity math.
I would choose active-passive write ownership because global ordering and RPO zero require one fenced serialization path.
- The passive region continuously receives synchronous or quorum-protected data, accepting WAN latency or reduced write availability during partition.
- Reads may run active-active if their freshness contract allows it, but only one region holds the current fencing token for writes.
- A timed promotion test must fence the old writer and restore writes inside ten minutes before the design meets its RTO claim.
Why interviewers ask this: A strong answer chooses a topology from ordering, data-loss, and recovery requirements rather than preferring active-active by default.
I would require sustained user-path failure from independent observers and proof that the destination is ready.
- At least three of five probes from separate networks must see over 5% errors or SLO-breaking latency for two minutes.
- The target region must have healthy dependencies, replication within RPO, and capacity for the full 40,000 requests per second.
- Failover uses hysteresis and a 15-minute minimum hold period so brief recovery cannot flap traffic between regions.
Why interviewers ask this: The interviewer evaluates failure detection, false-positive control, and destination readiness with explicit thresholds.
I would use asynchronous cross-region replication because synchronous acknowledgement cannot fit a 25 ms budget over an 80 ms link.
- Local writes acknowledge after durable regional commit, while replication lag is measured against the 30-second RPO.
- Alerts fire before lag consumes the full budget, and promotion is blocked when the replica position is outside it.
- Ledgers that truly require zero loss need a different latency contract or a quorum placement closer to writers.
Why interviewers ask this: A strong answer makes the latency-versus-data-loss decision explicit and avoids claiming incompatible guarantees.
I would require a consensus-backed lease and fence the former primary before the new region accepts writes.
- Promotion obtains a monotonic fencing token from an odd quorum placed across three independent failure domains.
- Storage and downstream consumers reject writes carrying an older token, even if the old primary is still reachable by some clients.
- Loss of quorum stops promotion and write availability; it never silently weakens the single-writer invariant to meet five minutes.
Why interviewers ask this: The interviewer checks consensus, fencing, and willingness to sacrifice write availability rather than corrupt ordered state.
I would run autonomous clusters per region and pre-provision the recovery region for the declared traffic shift.
- Each cluster has local ingress, DNS, registry, secrets, observability, and no control-plane dependency on the other region.
- The recovery region keeps enough warm capacity to reach 60,000 requests per second inside 15 minutes after cache warm-up.
- Global traffic moves only after data fencing and dependency checks, while GitOps keeps reviewed configuration equivalent without one shared runtime controller.
Why interviewers ask this: A strong answer preserves regional autonomy and includes capacity, data, and platform dependencies in the RTO.
I would create at least 20 repeatable cells and route each user deterministically to one cell.
- Every cell owns bounded compute, queues, caches, and data partitions so overload cannot consume a healthy cell's pool.
- Global identity and routing stay minimal and redundant because a shared dependency could erase the 100,000-user blast-radius goal.
- Capacity planning includes one spare cell and a controlled tenant-rebalancing path rather than moving users during an active failure.
Why interviewers ask this: The interviewer evaluates blast-radius arithmetic, dependency isolation, and the operating cost of cell recovery.
I would use a regional Vault performance replica with tested promotion and leases that outlive the 30-minute isolation window.
- Applications authenticate through regional workload identity and cache only current credentials, not reusable root or bootstrap tokens.
- Promotion includes seal or auto-unseal availability, storage health, token semantics, and fencing of the former active cluster.
- A game day proves all 80 services renew or degrade safely for 30 minutes without bypassing authorization.
Why interviewers ask this: A strong answer covers Vault replication, credential lifetime, promotion, and application behavior during isolation.
I would first challenge total global order and shard serialization by account if the invariant permits it.
- Per-account leaders preserve the order users need while independent shards process unrelated accounts in parallel.
- If total order is mandatory, one consensus leader and WAN quorum become the throughput and latency ceiling.
- A benchmark must sustain 50,000 writes per second through leader loss and show the exact write unavailability during quorum changes.
Why interviewers ask this: The interviewer checks whether the candidate narrows an expensive ordering guarantee and validates the unavoidable coordination cost.
I would run a full traffic and data exercise from external probes, not only stop application instances.
- The test removes regional ingress and the replication link, then measures detection, decision, promotion, traffic shift, and cache warm-up separately.
- Database positions and business checks prove data loss stays within 60 seconds and old-primary fencing prevents dual writes.
- The design passes only after two consecutive exercises restore user success within 12 minutes and failback is also rehearsed.
Why interviewers ask this: A strong answer validates end-to-end recovery and data correctness with repeated timed evidence.
The baseline forecast is about 57,800 requests per second, then I add failure and forecast headroom.
- The calculation is 25,000 multiplied by 1.15 to the sixth power, not a flat 90% addition.
- I benchmark sustainable throughput at the latency SLO and plan at least 30% uncertainty above the forecast.
- Zone-loss capacity is calculated separately, so normal placement cannot consume the headroom reserved for the largest tolerated failure.
Why interviewers ask this: The interviewer evaluates growth compounding, performance limits, uncertainty, and failure capacity as separate inputs.
I would provision each zone for 15,000 requests per second, giving 45,000 total and 30,000 after one zone fails.
- Normal peak uses about 67% of total capacity, which is the cost of immediate one-zone tolerance.
- The calculation includes database connections, queue throughput, egress, and cache capacity, not only application CPU.
- A zone-loss load test must hold latency and errors at 30,000 requests per second before the headroom is considered real.
Why interviewers ask this: A strong answer performs the failure-capacity math and applies it to every constrained dependency.
I would keep warm capacity for the 60-second burst because node autoscaling cannot beat a four-minute startup.
- Workload scaling uses a leading signal such as queue age or concurrency before CPU saturation.
- Placeholder capacity or a minimum node floor covers the measured burst while new nodes join.
- I separately measure cloud launch, bootstrap, CNI, image pull, and warm-up to reduce four minutes without pretending it is immediate.
Why interviewers ask this: The interviewer checks whether the candidate recognizes a timing mismatch and combines predictive scaling with reserved headroom.
I would size requests per workload from sustained percentiles and burst behavior, not the cluster-wide 35% average.
- CPU requests near p95 plus workload-specific margin improve placement, while latency-sensitive bursts may use uncapped spare CPU.
- Memory uses a higher working-set percentile because underestimation causes eviction or OOM rather than throttling.
- I canary new values and retain enough allocatable capacity for one-zone loss while checking throttling, Pending Pods, and SLOs.
Why interviewers ask this: A strong answer distinguishes averages, tail usage, CPU and memory behavior, and failure headroom.
I need at least 13 workers during recovery because 400 per second drains the burst while 100 per second handles new arrivals.
- The drain requirement is 240,000 divided by 600 seconds, then incoming traffic is added before dividing by worker throughput.
- I would provision about 16 workers for retries and processing variance, subject to broker and downstream limits.
- Queue age, not only depth, drives scaling and confirms the oldest message falls below ten minutes.
Why interviewers ask this: The interviewer checks queue-drain arithmetic and whether new arrivals, variability, and downstream capacity are included.
I would optimize and scale the primary vertically first because read replicas do not remove a write bottleneck.
- Query profiles, batching, indexes, connection overhead, and WAL or storage limits show whether CPU is the true constraint.
- A larger primary buys time with the least application complexity, but I document its hardware and failover ceiling.
- Sharding starts only when projected writes exceed that ceiling, using a key that spreads load and preserves required transaction boundaries.
Why interviewers ask this: A strong answer matches the scaling method to a measured write bottleneck and treats sharding as a deliberate complexity cost.
I would reproduce the production endpoint mix with an open workload model and test beyond both peak and failure capacity.
- Payload sizes, cache state, geographic latency, authentication, and background jobs match production rather than uniform empty requests.
- Load steps through 50%, 100%, and 130%, then repeats peak after removing one zone.
- The test records latency histograms, errors, saturation, queues, and dependency limits long enough to expose leaks and autoscaling delay.
Why interviewers ask this: The interviewer evaluates representative workload design, coordinated-omission avoidance, and explicit failure-capacity validation.
I would plan regional reservation ranges six weeks ahead and maintain a portable overflow tier for forecast error.
- Demand is split into committed baseline, launch-driven upside, and batch work that can move in time or region.
- Each region tracks quota, reserved GPUs, allocatable GPUs, queue age, and the lead time to add the next block.
- I reserve the critical baseline plus measured uncertainty, while interruptible batch absorbs the 40% range instead of duplicating it everywhere.
Why interviewers ask this: A strong answer combines long provisioning lead time, uncertain demand, regional constraints, and movable workloads.
I would measure successful eligible checkouts and completion below a user-relevant latency threshold at the transaction boundary.
- Availability counts committed orders as good and 5xx, invalid responses, or lost confirmations as bad.
- Latency counts eligible journeys completed within 800 ms, while fast errors remain bad outcomes.
- Dependency metrics explain failures, but retries and internal calls do not inflate the 10,000-request-per-second user denominator.
Why interviewers ask this: The interviewer checks precise good and valid events, latency semantics, and measurement at the user boundary.
I would assign the strictest target to the class with the highest user and correctness cost, not the highest volume.
- Interactive writes may receive 99.99% availability with a 500 ms latency threshold.
- Interactive reads can use 99.9% if caching or stale fallback preserves value, while batch completion may use 99% within a deadline.
- Every request maps to one class so high-volume reads cannot hide failure of low-volume critical writes.
Why interviewers ask this: A strong answer segments SLOs by business impact and prevents aggregate traffic from masking critical paths.
Locked questions
- 21
An API depends on three services each claiming 99.9%. Why is multiplying them into 99.7% not a sufficient service SLO design?
designsloapi - 22
A streaming endpoint legitimately runs for 30 seconds, while normal API calls should finish in 400 ms. How would the latency SLI handle both?
endpointslatencystreaming - 23
Calculate the monthly error budget for 99.9% availability at 100 million valid requests.
reliabilityavailabilityerror-budget - 24
Design multi-window burn-rate alerts for a 99.9% SLO over a 30-day window.
designsloalerting - 25
A critical admin API receives only 200 requests per day. How would you define a meaningful 99.9% reliability objective?
reliabilityapi - 26
Search returns cached results up to five minutes old during overload. Should those requests count as good in a 99.9% SLO?
slocaching - 27
A checkout journey has a 99.95% SLO and depends on payment, inventory, and identity. How would you set internal objectives?
slo - 28
Design OpenTelemetry collection for 500 services across three regions at 1 million telemetry events per second.
designobservability - 29
How would you operate Prometheus for 30 clusters and 15 million active series with 13 months of retention?
retentionmonitoring - 30
Metric volume grows from 4 million to 20 million series after teams add user_id and raw URL labels. What controls do you design?
designmonitoring - 31
200 services emit metrics, logs, and traces, but a release cannot be followed across all three. What correlation contract would you set?
correlationmonitoring - 32
A tracing platform receives 2 million spans per second but can retain only 5%. How would you sample?
- 33
Design observability to remain useful when traffic reaches 3x normal and telemetry volume threatens the collectors.
designobservability - 34
A platform has 800 services and currently sends 600 pages per day. What alerting design would you implement?
alertingdesign - 35
Grafana has 200 dashboards, and core incident views take 18 seconds to load. How would you redesign them?
incidentsincident-managementmonitoring - 36
Define severity levels for a platform serving 80 services and 5 million users.
severity-priority - 37
Design a sustainable 24/7 on-call rotation for ten SREs supporting 60 services.
on-calldesign - 38
A sev1 bridge attracts 30 engineers in ten minutes. What incident-command structure should already exist?
incidentsincident-management - 39
How would you design follow-the-sun incident handoff across three regions during an eight-hour sev1?
incidentsincident-managementdesign - 40
A postmortem program creates 120 actions per quarter but closes only 25%. How would you redesign the process?
incidentspostmortemsconcurrency - 41
A payment provider may be unavailable for 20 minutes, while your API must answer within two seconds. What degraded mode do you design?
designapi - 42
A multi-region API serves 100,000 requests per second, while one tenant is limited to 1,000. How would you divide enforcement across regions?
api - 43
A service is safe at 60,000 requests per second but may receive 90,000. How would you shed load?
- 44
A four-stage pipeline receives 8,000 events per second, but stage three can process only 6,000. How do you design backpressure?
designresiliencebackpressure - 45
A service makes 50,000 dependency calls per second. How would you cap retries at 10% extra load?
dependencies - 46
Set circuit-breaker behavior for a dependency that normally handles 20,000 calls per second at 0.2% errors.
dependencies - 47
How would you validate load shedding, retry budgets, and circuit breakers together at 80,000 requests per second?
resiliencevalidationload-management - 48
Design the first chaos experiment for a three-zone service with 99.95% availability and no prior game days.
designavailabilitychaos-engineering - 49
Plan a regional game day for RTO 15 minutes and RPO one minute without risking all customers.
disaster-recovery - 50
Forty tier-1 services need recurring chaos coverage. What program would you run over 12 months?
chaos-engineeringcoverage - 51
At 09:02 UTC, checkout 5xx jumps from 0.2% to 18% two minutes after release 4.17 reaches 40% of Pods; what do you do in the first 10 minutes?
- 52
At 14:10 UTC, a payment dependency slows from 80 ms to 3 seconds and three client layers amplify 8,000 requests per second into 54,000; how do you contain the retry storm?
resiliencedependencies - 53
Between 16:00 and 16:12, API p50 remains 45 ms but p99 rises from 320 ms to 4.8 seconds for 7% of users in one zone; how do you diagnose it?
api - 54
At 11:25 UTC, PostgreSQL reaches 500 of 500 connections, API errors hit 22%, and CPU is only 38%; what is your incident decision?
incidentsincident-managementapi - 55
At 08:40 UTC, Kubernetes cannot create nodes because the cloud account has used 2,000 of 2,000 vCPUs while 180 Pods remain Pending; how do you recover?
kubernetes - 56
At 03:15 UTC, an Elasticsearch data node is 97% full, watermarks block writes, and log ingestion is 140 GB per hour; what do you do before disk reaches 100%?
search - 57
At 19:00 UTC, Kafka consumer lag grows from 20,000 to 9 million records after one downstream shard slows to 400 events per second; incoming traffic is 6,000 per second. What do you do?
shardingkafka - 58
At 12:01 UTC, a DNS change lowers successful lookups from 99.99% to 91% for Android clients with a 30-minute TTL; what is your recovery plan?
recoverydns - 59
At 06:50 UTC, 28% of clients reject a renewed TLS certificate because one intermediate chain is missing, and the old certificate expires in 40 minutes; what do you do?
tls - 60
At 02:20 UTC, Redis fails over in 18 seconds, but 600 application Pods reconnect simultaneously and drive CPU to 96% with 35% timeouts; how do you stabilize it?
redisresilience - 61
A release at 10:00 burns 18% of a 30-day error budget in 25 minutes, but only a new reporting feature is failing; do you roll back the whole release?
error-budgetreliabilityrollback - 62
At 17:30 UTC, enabling a recommendation flag for 50% of users raises checkout p99 from 700 ms to 1.9 seconds without increasing errors; what do you decide?
- 63
At 13:05 UTC, a schema migration has held an ACCESS EXCLUSIVE lock for 95 seconds, 1,800 writes are queued, and replication lag is 70 seconds; what do you do?
schemamigrationsreplication - 64
At 09:45 UTC, a 2 TB cache expires at once and origin traffic rises from 15,000 to 110,000 requests per second; database CPU reaches 89%. How do you stop the stampede?
zero-to-onedatabasecaching - 65
At 18:20 UTC, a shipping provider starts returning 429 to 70% of calls, while 24,000 orders are waiting and its quota resets in 35 minutes; what do you do?
- 66
At 04:12 UTC, a Kubernetes node upgrade evicts 220 Pods, 65 remain Pending, and API availability falls to 96%; what is your response?
availabilityapikubernetes - 67
From 01:00 to 05:00, service memory grows 120 MB per hour per Pod; 14 of 80 Pods have OOM-killed and p99 is 1.4 seconds. How do you manage the leak overnight?
memory - 68
At 15:40 UTC, a Java service shows 9-second stop-the-world pauses every 2 minutes after heap occupancy reaches 82%, but CPU is 45%; what do you do?
data-structures - 69
At 20:05 UTC, one availability zone shows 4% packet loss, gRPC retries triple traffic to 72,000 calls per second, and the provider has no ETA; what do you do?
availabilitygrpc - 70
At 10:30 UTC, 12 of 100 API Pods receive 46% of traffic and hit 95% CPU while the rest stay below 30%; how do you handle the imbalance?
restsoft-skills - 71
At 21:15 UTC, one tenant sends 18,000 requests per second against a contracted 2,000, raising shared API p99 to 2.2 seconds for 300 tenants; what do you do?
api - 72
At 07:10 UTC, NTP drift reaches 95 seconds on 32 nodes, causing 14% of OAuth tokens to appear expired; how do you recover without accepting invalid tokens?
oauthtokensiac - 73
At 05:55 UTC, the identity provider rotates signing keys, 38% of API requests fail JWT validation, and the JWKS endpoint is timing out for 6 seconds; what do you do?
jwtvalidationendpoints - 74
At 09:20 UTC, a release adds customer_id to Prometheus labels, active series jump from 6 million to 48 million, and collectors start dropping 30% of samples; what do you do?
monitoring - 75
At 23:40 UTC, Grafana and tracing are unavailable during a 12% API error spike, but application logs and cloud metrics remain; how do you lead diagnosis for 20 minutes?
monitoringapi - 76
After a 25-minute outage, 4.5 million jobs are queued, normal capacity is 8,000 per second, and new arrivals are 6,000 per second; how do you recover without causing a second outage?
capacitydata-structurescapacity-planning - 77
At 18:00 UTC, traffic reaches 140,000 requests per second against a tested safe limit of 100,000, and checkout errors climb to 9%; which load do you shed?
- 78
At 10:05 UTC, mobile clients retry every 100 ms after 503 responses, turning 20,000 user actions per second into 160,000 requests; no client update can ship for 24 hours. What do you do?
resilience - 79
At 03:00 UTC, the primary region is down, the replica is 75 seconds behind, and 11,000 orders may be missing; the business asks for immediate failover. What do you decide?
replicationfailover - 80
At 13:40 UTC, reconciliation finds 620 duplicate charges created during a 17-minute timeout incident; service is now healthy. What do you do next?
incidentsincident-managementreact - 81
A 42-minute outage began at 09:07 when an engineer changed an Envoy timeout from 2 seconds to 200 ms; rollback took 31 minutes because ownership was unclear. How do you run the postmortem?
ownershipincidentspostmortems - 82
The same queue overflow caused 3 outages in 6 months despite a completed runbook action; the latest impact lasted 28 minutes. What do you change after review?
runbooksdata-structures - 83
Four SREs spend 26 hours per week manually restarting failed batch jobs; each automation attempt would take about 120 engineering hours. Do you automate now?
batch - 84
Two engineers spend 9 hours per week renewing 240 certificates, and one missed renewal caused a 16-minute outage; what automation do you approve?
- 85
The on-call SRE spends 14 hours per week granting temporary production access, with 60% of requests arriving after hours; what do you change?
on-call - 86
A release engineer spends 18 hours per week running 45 manual deployment steps, and 7% of releases need correction; what is your automation sequence?
deployment - 87
A team spends 11 hours per week investigating 320 alerts, but only 24 require action; what do you do during the next 2 weeks?
alerting - 88
During a 12-hour shift, the primary receives 17 pages, including 6 after midnight, and misses one 10-minute acknowledgement target; what do you do before the next shift?
- 89
At 02:10 UTC, one on-call engineer receives a database Sev1 and a separate CDN Sev1 within 3 minutes; only 4 responders are available. How do you allocate them?
on-calldatabase - 90
A Sev1 has run for 7 hours; the outgoing team has 25 minutes to hand over, with 3 mitigations active and replication lag at 48 seconds. How do you transfer command?
replication - 91
A junior on-call engineer sees errors rise from 0.3% to 6% after a 10% canary but wants 20 more minutes of diagnosis before rollback; how do you mentor in the moment?
mentoringon-callrollback - 92
A mid-level SRE manually raises a PostgreSQL connection limit from 400 to 800 during 2 incidents, worsening both; how do you coach and control the next incident?
incidentsincident-managementpostgres - 93
A new incident commander spends 15 minutes debugging one Pod while 12 responders receive no assignments during a 9% outage; how do you intervene and mentor?
mentoringincidentsincident-management - 94
Traffic grows 12% weekly for 5 weeks, database storage has 9 days left, and adding a shard takes 14 days; what action do you take today?
databasesharding - 95
During a traffic spike, finance asks you to remove 30% reserved compute to save $18,000 per day while one-zone headroom is only 12%; what do you decide?
- 96
A rollback from version 7.4 to 7.3 lowers errors from 14% to 4% but then stalls because 7.3 cannot read 8% of newly written records; what do you do?
rollback - 97
At 22:00 UTC, an L7 attack raises traffic from 25,000 to 900,000 requests per second, WAF CPU is 88%, and 3% of legitimate logins fail; how do you respond?
- 98
At 01:30 UTC, ransomware indicators appear on 6 build agents while a production hotfix is due in 20 minutes; security isolates CI. How do you keep the outage response safe?
- 99
At 11:11 UTC, an inventory call slows to 1.5 seconds, checkout has a 2-second deadline, and 4 sequential retries push p99 to 7 seconds; what do you change during the incident?
incidentsestimationincident-management - 100
After a 63-minute outage, errors have been below 0.2% for 8 minutes, but 1.2 million messages remain queued and replication lag is 35 seconds; when do you declare recovery?
replicationdata-structures