Skip to content

GCP Engineer interview questions

100 real questions with model answers and explanations for Staff Cloud Engineer candidates.

See a GCP Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

designslo

I would build the landing zone around failure, ownership, and residency boundaries rather than mirror the org chart blindly.

  • I would place bootstrap, security, networking, and logging projects in dedicated platform folders, then split workload folders by production status and US or EU residency.
  • A project factory would create projects, APIs, budgets, IAM groups, Shared VPC attachments, log sinks, and baseline policies from one approved request in under 30 minutes.
  • Folder-level Organization Policy and IAM would supply defaults, while project-level grants stay narrow so one team cannot administer another team's 220-project estate.
  • I would run the factory and policy checks from redundant regional runners, measure provisioning success and age, and keep a documented manual path for the 99.95% SLO.

Why interviewers ask this: The interviewer is evaluating whether you can turn hierarchy, residency, networking, and automation into an operable enterprise landing zone.

designslonetworking

I would use several bounded Shared VPC domains instead of making one host project a company-wide blast radius.

  • Separate production, nonproduction, and regulated host projects would own regional subnets, Cloud Routers, NAT, firewall policy, and connectivity, with two platform teams required for changes.
  • Subnet-level IAM would let each service-project group consume only assigned primary and secondary ranges, and a central IPAM service would prevent overlap across all 90 projects.
  • I would reserve non-RFC1918 privately used ranges or translate the colliding 10.0.0.0/8 routes at the hybrid boundary until on-premises renumbering finishes.
  • Regional paths, redundant Cloud Routers, flow logs, quota alarms, and synthetic reachability probes would make the 99.99% SLO measurable rather than assumed.

Why interviewers ask this: A strong answer limits Shared VPC blast radius and resolves address overlap without weakening ownership or availability.

designcapacityslo

I would use a cell-based model with pooled tenants by default and dedicated cells for the 60 tenants that require a hard boundary.

  • Each regional pooled cell would cap at roughly 500 tenants and have its own GKE or Cloud Run compute, data partition, service accounts, quotas, and failure budget.
  • Regulated tenants would receive separate projects and data stores when contractual isolation requires it, while the same platform APIs keep operations consistent.
  • Tenant identity would be carried in signed claims and enforced again at the data layer, with per-tenant rate limits preventing any tenant from exceeding 2% of cell capacity.
  • Global placement metadata would route tenants to their home region, and tested cell evacuation plus bounded replication would support the 99.95% SLO without globalizing tenant data.

Why interviewers ask this: The interviewer wants a concrete isolation model that balances hard boundaries, noisy-neighbor control, regional placement, and operating cost.

configerror-handling

I would make centrally inherited preventive controls the default and treat every exception as expiring code.

  • Organization Policy would block service account keys, unapproved regions, public IPs, and unrestricted sharing at the organization or folder level, using custom constraints where managed constraints are insufficient.
  • Terraform plan checks and policy tests would fail before apply, while Cloud Asset Inventory queries would detect resources created outside the approved path within the one-hour control window.
  • An exception workflow would require an owner, risk, compensating control, and expiry, then place the workload in a narrowly scoped folder or tag-based rule for no more than 30 days.
  • Policy bundles would pass tests in a staging folder before promotion, and project-factory templates would keep compliant delivery comfortably inside the 48-hour target.

Why interviewers ask this: The interviewer is checking whether prevention, detection, safe rollout, and time-bounded exceptions form one governance system.

design

I would make billing ownership part of project creation and calculate chargeback from the detailed billing export.

  • Every project would carry validated cost-center, product, environment, and owner metadata, while the resource hierarchy supplies a fallback when a service cannot preserve labels.
  • Detailed billing export to BigQuery would allocate direct spend first, then distribute Shared VPC, observability, and support costs by measured usage such as bytes, vCPU-hours, or active services.
  • Reservations and committed use discounts would be amortized back to beneficiaries instead of credited only to the billing account, keeping unit costs economically honest.
  • Daily dashboards would track unallocated spend, team budgets, forecasts, and the 12% platform ratio, with day-one alerts leaving two days to resolve mapping gaps.

Why interviewers ask this: A strong answer distinguishes attribution from fair allocation and handles discounts and shared services explicitly.

designslolatency

I would put one global external Application Load Balancer in front of regional origins and make cache behavior part of the capacity design.

  • Cloud CDN would cache versioned static assets and safe API responses with explicit cache keys and TTLs, targeting an origin offload ratio that keeps the $120,000 budget viable.
  • Cloud Armor would apply managed WAF rules, adaptive protection, rate limits, and bot controls before traffic reaches the four origins.
  • Regional backends would expose health checks and enough warm capacity for one-region loss, while weighted routing would support canaries without a DNS cutover.
  • Edge, cache, Armor, and origin latency would have separate SLIs, and synthetic probes from key markets would enforce the 250 ms p95 and 99.99% objectives.

Why interviewers ask this: The interviewer is evaluating how load balancing, caching, security, failover capacity, and cost fit one global traffic design.

designlatency

I would use Dedicated Cloud Interconnect as the primary path and HA VPN only as a lower-capacity emergency path.

  • To qualify for the 99.99% topology, I would provision four Interconnect connections across two metropolitan areas, place the two connections in each metro in different edge availability domains, and create one VLAN attachment on every connection.
  • Separate Cloud Routers, customer routers, power, and carrier paths would remove shared failure points, while BGP priorities and tested failure detection target recovery under five minutes.
  • HA VPN would preserve control and critical traffic during a broader Interconnect failure, but it would not be sized or represented as carrying the full steady 20 Gbps.
  • Each attachment would normally remain below 50% utilization, and loss, latency, learned routes, and four-path failover tests would prove capacity and availability.

Why interviewers ask this: A strong answer earns the availability target with physical diversity and states the real capacity limit of the VPN fallback.

communication

I would use NCC's global hub model and make trust boundaries explicit rather than invent regional hubs.

  • An NCC hub is global; I would use separate hubs for trust domains that must never exchange routes, and route tables plus spoke groups to segment production, nonproduction, and regulated connectivity within a hub.
  • The regional components are Cloud Routers, VLAN attachments, HA VPN tunnels, and regional hybrid spokes, while VPC spokes connect networks to the global control plane.
  • Import and export rules would let app spokes reach shared inspection, DNS, and approved on-premises prefixes without learning other app-spoke routes.
  • Redundant regional attachments, summarized advertisements, route-diff tests, Connectivity Tests, and convergence probes would verify the two-minute objective.

Why interviewers ask this: The interviewer checks whether NCC routing domains and route exchange enforce segmentation rather than merely connect every network.

designslonetworking

I would publish one regional service attachment per serving region and give each consumer a private endpoint in its own VPC.

  • Regional producer backends behind supported internal load balancers would scale independently, so a consumer never receives routes to the producer VPC.
  • Connection preferences, project allowlists, endpoint quotas, and dedicated NAT subnets would control which of the 40 projects can attach and make each consumer visible.
  • Private DNS would return the endpoint appropriate to the consumer region, while clients would retry another published region only within the payment API's idempotency contract.
  • Per-endpoint connection count, rejection, latency, and regional capacity would be measured against p99 100 ms and provisioned for one-region loss.

Why interviewers ask this: The interviewer is evaluating whether you understand PSC's regional publishing model, consumer isolation, DNS, and failover implications.

designdnsslo

I would centralize hybrid resolution paths while keeping authoritative ownership and split-horizon behavior explicit.

  • Cloud DNS private zones would be associated only with authorized VPCs, and DNS peering would expose central service zones without duplicating 300 zone copies.
  • Inbound forwarding endpoints would serve on-premises queries, while outbound forwarding policies would send corporate suffixes to redundant resolvers in both data centers.
  • Conditional rules would have one owner and no circular forwarding path, with response policies reserved for controlled overrides rather than ad hoc records.
  • Query logs, synthetic lookups, resolver health, and a tested secondary forwarding route would enforce the 99.99% SLO and five-minute recovery target.

Why interviewers ask this: A strong answer separates authority, visibility, forwarding direction, and failure testing in a large hybrid DNS design.

formsdesign

I would keep private Google API traffic off the internet and apply egress controls by protocol rather than treat Cloud NAT as a filter.

  • Private Google Access or Private Service Connect endpoints would reach supported Google APIs, Secure Web Proxy would enforce the 500-domain HTTP and HTTPS policy, and Cloud NGFW or approved appliances would inspect required non-web flows.
  • Cloud NAT is a regional, distributed managed data plane with no zonal gateway to shard and no fixed gateway-throughput ceiling; VM and network-interface bandwidth still bound the 25 Gbps path.
  • I would size reserved external IPs and minimum or maximum ports per VM from peak concurrent endpoints, connection-opening rate, destination fan-out, protocol timeouts, endpoint-independent mapping needs, and regional address and NAT quotas.
  • NAT, proxy, firewall, and flow telemetry would track allocation failures, ports per VM, connection rate, bytes, blocked flows, and inspection cost against the $80,000 budget.

Why interviewers ask this: The interviewer wants protocol-aware egress controls, private API access, scalable NAT, and a credible cost model.

deploymentkubernetesdesign

I would run an independent regional cluster in each region and add explicit fleet services rather than assume registration supplies them.

  • Fleet Workload Identity would establish fleet-wide workload identity, while Config Sync and Policy Controller would distribute configuration and enforce policy across the three clusters.
  • Multi Cluster Services would export service discovery and Multi Cluster Gateway would program the global traffic surface; fleet registration by itself provides none of those runtime guarantees.
  • Services would use the same immutable artifact and contract, with region-specific values limited to endpoints, capacity, and residency, and each surviving region would hold tested failover headroom.
  • Accepted writes would commit to a datastore meeting RPO zero before success, and quarterly regional withdrawal tests would verify Multi Cluster Gateway health and the ten-minute RTO independently of Kubernetes repair.

Why interviewers ask this: The interviewer is checking whether the fleet design separates stateless regional compute from the data guarantees behind RTO and RPO.

kubernetesslo

I would offer Autopilot as the default workload tier and a smaller Standard tier for the six privileged-agent exceptions.

  • Autopilot removes node administration and bills requested pod resources, which fits the 70 stateless services if requests, concurrency, and startup time are tuned.
  • Standard clusters would host workloads that truly require privileged node access, custom node configuration, or capacity controls unavailable in the default Autopilot operating model.
  • Both tiers would share artifact, identity, policy, observability, and deployment interfaces so teams do not build two unrelated platforms.
  • I would compare cost per request and SLO attainment monthly, because oversized Autopilot requests or idle Standard nodes can each break the $90,000 target.

Why interviewers ask this: A strong answer makes a workload-based platform choice and preserves one developer contract across two operating modes.

kubernetesiacdecision-making

I would adopt GKE Enterprise only if its fleet governance replaces enough separate tooling and labor to justify the $40,000 cap.

  • Fleet membership, Config Sync, Policy Controller, and Cloud Service Mesh would be evaluated against the 16-cluster configuration and policy use cases, not purchased as a label.
  • I would run a two-cluster pilot that measures configuration convergence under 15 minutes, policy coverage, upgrade effort, and on-premises connectivity failure behavior.
  • Workload portability would remain based on Kubernetes APIs and Git, while Google-specific fleet features would be isolated behind documented platform contracts.
  • The decision gate would compare license and operations cost with the existing stack, and I would decline the package if governance SLO or net cost does not improve.

Why interviewers ask this: The interviewer is evaluating current GKE Enterprise capabilities as a measurable platform choice rather than using the former Anthos name as a strategy.

designservice-meshslo

I would use managed Cloud Service Mesh with a narrow traffic-policy baseline and prove its data-plane cost before broad rollout.

  • Fleet-wide workload identity and strict mTLS would replace shared certificates, with authorization policies tied to service identities rather than pod IPs.
  • Timeouts, retries, and circuit breaking would be set per dependency, with retry budgets preventing a failing service from multiplying 60,000 requests per second.
  • Sidecar CPU, memory, connection pools, and telemetry sampling would be load-tested so mesh overhead stays within the 15 ms p99 allowance.
  • I would canary mesh enrollment by namespace and retain direct rollback, while control-plane and proxy health become explicit components of the 99.99% SLO.

Why interviewers ask this: A strong answer treats identity, resilience policy, data-plane overhead, and rollout safety as measurable mesh design concerns.

designslodeployment-strategies

I would use GKE multi-cluster Gateway to program one global load-balancing surface from fleet-aware Gateway API resources.

  • Multi-cluster Services would expose identical backends from all three clusters, with health checks and service capacity defined independently per region.
  • HTTPRoute weights would send 5% to the canary backend, and immutable releases would allow an immediate route-weight rollback without changing DNS.
  • Removing session affinity keeps requests portable, while applications externalize session state and use idempotency for retries during regional shifts.
  • Each pair of surviving regions would hold failover headroom, and synthetic probes plus scheduled region withdrawal would verify the five-minute and 99.99% targets.

Why interviewers ask this: The interviewer checks whether Gateway API, backend health, statelessness, canary routing, and capacity make multi-cluster failover real.

kubernetesdesignonboarding

I would use namespaces for ordinary teams and dedicated clusters or node boundaries where the PCI threat model requires stronger isolation.

  • The onboarding workflow would create namespace, Workload Identity Federation bindings, RBAC groups, NetworkPolicies, quotas, limits, and observability ownership within one day.
  • ResourceQuota and admission policy would cap each shared tenant at 5% of allocatable CPU unless an approved capacity request changes the limit.
  • Default-deny east-west networking and service identity would prevent namespace membership from becoming implicit trust between teams.
  • PCI workloads would use separate projects and clusters when kernel, node, or administrator isolation is required, accepting the extra cost rather than overstating namespace security.

Why interviewers ask this: A strong answer matches the tenancy boundary to the threat model and automates quotas, identity, networking, and onboarding.

designcapacityslo

I would separate price coverage from capacity assurance and keep a reliable on-demand base under the Spot layer.

  • Release channels, maintenance windows, version-skew checks, and a canary cluster would promote upgrades through nonproduction and production inside 30 days.
  • Regional clusters, topology spread, PodDisruptionBudgets, and controlled surge upgrades would preserve replicas while a zone or node pool drains.
  • Eligible baseline usage would receive committed use discounts to reduce price, but zonal reservations sized for critical node pools would supply capacity after one-zone loss; a CUD alone reserves no capacity.
  • Spot pools would run only checkpointable or overprovisioned work, and rightsizing, reservation coverage, disruption tests, and unit cost would prove 35% savings without breaching 99.95%.

Why interviewers ask this: The interviewer is evaluating whether upgrade velocity, zonal resilience, capacity guarantees, and Spot economics are designed together.

designsloserverless-containers

I would standardize a regional Cloud Run service contract and place multi-region services behind a global external Application Load Balancer.

  • The template would set service identity, ingress, secrets, CPU, memory, concurrency, maximum instances, logging, and deploy policy instead of exposing every knob to 30 teams.
  • Latency-sensitive services would keep measured minimum instances, while development and tolerant services scale to zero to protect the $70,000 budget.
  • Maximum instances and downstream-aware concurrency would stop a 60,000-request-per-second surge from exhausting databases or regional quotas.
  • Per-revision latency, cold starts, saturation, cost per request, and regional health would drive gradual rollout and the 99.95% SLO.

Why interviewers ask this: A strong answer turns Cloud Run settings, regional routing, downstream protection, and unit cost into a reusable platform.

designserverless-containersevents

I would normalize every source as a CloudEvent and keep the invoice effect idempotent beyond the trigger's delivery semantics.

  • Eventarc filters would route narrow event types to small Cloud Run functions or services with dedicated identities, avoiding one universal handler.
  • Critical flows needing controlled retention, dead letters, or replay would pass through an explicit Pub/Sub topic and subscription rather than depend only on a direct trigger.
  • The handler would claim the immutable event ID in a transactional invoice ledger before creating the invoice, so retries cannot repeat the business effect.
  • Event age, retry count, dead-letter volume, and end-to-end p95 would be monitored, with 24-hour retention and replay tested against schema versions.

Why interviewers ask this: The interviewer checks whether Eventarc routing, function scope, durable retries, and business idempotency are separated correctly.

Locked questions

  • 21

    Design Pub/Sub backpressure for 50,000 messages per second with ten-minute bursts to 300,000, 1 KB messages, a five-minute freshness SLO, and consumers that can scale to 400,000 messages per second within two minutes.

    designslomessaging
  • 22

    Design order processing on Pub/Sub for 100,000 events per second and two million tenants, with per-order sequencing, zero duplicate charges, consumers in two regions, and recovery within 15 minutes.

    designmessagingconcurrency
  • 23

    Design an asynchronous checkout API sustaining 20,000 requests per second, returning within 200 ms, dispatching to a fulfillment fleet capped at 25,000 operations per second, completing within 15 minutes, and preserving fairness across 500 tenants.

    designapiasync
  • 24

    Design batch execution for 20,000 container jobs per day, where 60% finish within 15 minutes, 10% require GPUs, all finish within six hours, and monthly compute spend is capped at $25,000.

    designbatchcontainers
  • 25

    Choose Cloud SQL, AlloyDB, or Spanner for a relational SaaS control plane doing 25,000 transactions per second in three regions, with p99 reads below 50 ms, RPO zero on region loss, a 99.999% SLO, and 40 TB of data.

    sqltransactionsslo
  • 26

    Design a Spanner ledger for 8,000 writes per second from five regions, externally consistent balances, write p99 below 120 ms, RPO zero, and a 99.999% availability target.

    design
  • 27

    Design a BigQuery lakehouse for 3 PB of existing data, 50 TB of daily ingestion, 200 analysts, dashboard p95 below five seconds, EU data residency, and a $180,000 monthly analytics budget.

    bigquerydesign
  • 28

    Design a Dataflow pipeline for one million events per second, ten-second p99 freshness, 15 minutes of allowed lateness, seven-day replay, RTO of 15 minutes, and no duplicate business records.

    designci-cd
  • 29

    Design Cloud Storage for 5 PB of media in two regions, with 80% of objects untouched after 90 days, retrieval within one hour, at least eleven nines of durability, RPO of 15 minutes, and a $60,000 monthly storage budget.

    designobject-storage
  • 30

    Design data residency for 8,000 SaaS tenants split between the EU and US, including 200 regulated EU tenants, with zero customer-data crossing, two serving regions per geography, and a 99.95% SLO.

    designslocloud
  • 31

    Design backup for 500 TB across databases, GKE persistent data, and Cloud Storage, with RPO of 15 minutes, RTO of four hours, 35-day normal retention, 90-day immutable copies, and a $40,000 monthly budget.

    designbackupsobject-storage
  • 32

    Design disaster recovery for a 40 TB AlloyDB application doing 30,000 transactions per second, with RPO below one minute, RTO of 15 minutes, a 99.99% SLO, and DR cost no more than 45% of primary cost.

    designslotransactions
  • 33

    Design zero-trust access for 6,000 employees, 800 workloads, and 250 projects, with no service account keys, privileged sessions limited to one hour, and access revocation within 15 minutes.

    sessionszero-trustidentity
  • 34

    Design VPC Service Controls for 120 projects holding 4 PB in BigQuery and Cloud Storage, with three vendors, zero unauthorized exfiltration, vendor onboarding within one hour, and uninterrupted CI pipelines.

    designnetworkingobject-storage
  • 35

    Design CMEK for 300 PCI projects in two regions, with key rotation every 90 days, emergency revocation within 15 minutes, a 99.99% SLO, and no application access to raw key material.

    designslo
  • 36

    Design Access Context Manager device-context governance for 6,000 employees and 300 contractors in 25 countries, requiring managed devices for production access, revocation within 15 minutes, contractor exceptions below eight hours, and a 99.9% access SLO.

    designsloerror-handling
  • 37

    Design Security Command Center for 300 projects and one million assets in PCI scope, routing supported event-driven critical findings within five minutes, containing high-confidence threats within 15 minutes, and closing owned findings within 24 hours.

    eventsdesign
  • 38

    Design a Binary Authorization and SLSA-aligned supply chain for 200 services and 100 daily deployments, admitting only provenance-backed builds, quarantining running images with new critical findings within 15 minutes, and supporting audited emergency deployment.

    authsupply-chaindesign
  • 39

    Design audit evidence for SOC 2 and PCI across 180 projects in EU and US regions, with 12-month retention, query results within ten minutes, no log gaps, and evidence delivery within one business day.

    retentionqueriesdesign
  • 40

    Design an infrastructure platform using Terraform and Crossplane for 70 teams, 500 GCP projects, 30 resource requests per day, self-service delivery within 20 minutes, and control-plane RTO of two hours.

    terraformiacdesign
  • 41

    Design Terraform state for 1,000 resources across 80 environments and 20 engineers, with concurrent plans, recovery within 30 minutes, and any failed apply limited to one application environment.

    terraformdesignconcurrency
  • 42

    Design policy as code for 400 Terraform modules and 40 teams, with 95% compliant delivery, critical violations blocked in under five minutes, and exceptions expiring after 14 days.

    terraformdesignerror-handling
  • 43

    Design GitOps for 15 GKE clusters and 300 services, with 200 deployments per day, convergence within ten minutes, rollback within five minutes, and no human production credentials.

    deploymentrollbackgitops
  • 44

    Design a Backstage self-service portal for 600 developers and 40 teams, with a new production-ready service created within 15 minutes, 80% golden-path adoption in six months, and a 99.9% portal SLO.

    designslodecision-making
  • 45

    During week 6 of migrating 450 discontinued Deployment Manager stacks across 120 projects, a Terraform import wave proposes recreating 38 production resources; how do you recover without downtime and keep recreation below 1%?

    terraformdeployment
  • 46

    Design an observability platform for 120 services and 40 teams, with 99.95% availability SLOs, p99 latency objectives of 300 ms, burn alerts within five minutes, and a $60,000 monthly telemetry budget.

    designsloobservability
  • 47

    Design a quarterly multi-region DR game day for a commerce platform with 60 services, 30 TB of data, 80,000 requests per second, a 99.99% SLO, RTO of 20 minutes, and RPO of five minutes.

    designslo
  • 48

    Reduce a $4.2 million annual GCP budget by 20% within three months across 45 teams, while preserving a 99.95% SLO and keeping peak capacity 30% above forecast.

    capacityslo
  • 49

    Design migration waves for 25 legacy applications on 180 VMs with 120 TB of data, two data centers closing in nine months, four hours maximum downtime per application, RPO of 15 minutes, and a 99.9% post-migration SLO.

    designslomigrations
  • 50

    Set a platform standard for 25 teams and 140 services split between GKE and Cloud Run, with 70% adoption in six months, onboarding within one day, a 99.95% SLO, and exceptions below 10%.

    onboardingdecision-makingslo
  • 51

    A new Organization Policy causes 83% of a rollout from 10 canary projects to 200 production projects to fail, while project provisioning has a 30-minute SLO and the launch starts in four hours; how do you recover?

    slodeployment-strategiesguardrails
  • 52

    Cloud Asset Inventory reports 27 unexpected Owner-level grants across 40 projects in 18 minutes against a baseline of zero per day, and the IAM response SLO is 15 minutes; what do you do?

    slo
  • 53

    A service-account key with access to six production projects has been public for 14 minutes, outbound traffic is 3.5 times its normal 2 Gbps baseline, and key containment has a 10-minute SLO; how do you respond?

    slo
  • 54

    An Access Context Manager change denies 72% of IAP and Cloud Console requests from 4,800 managed laptops across 25 countries, versus a 99.9% workforce-access SLO, and a payroll cutover starts in two hours; how do you distinguish stale device context from real noncompliance and recover?

    slocontext-managers
  • 55

    Disabling one Cloud KMS key version causes 78% of requests from 60 services to fail within six minutes, versus a 0.005% error baseline and a 99.99% serving SLO; how do you restore service?

    slo
  • 56

    A Certificate Authority Service-backed renewal controller stalls for six hours; 2,400 of 12,000 workload mTLS certificates expire within 30 minutes, handshake failures reach 14% from a 0.005% baseline, and the serving SLO is 99.99%; how do you prevent a fleet outage?

    mtlsslo
  • 57

    One Interconnect BGP session advertises an unintended 10.42.16.0/20 at priority 80, displacing the expected 10.42.16.0/20 at priority 100 and blackholing 30% of 12 Gbps hybrid traffic, while recovery must take under five minutes; how do you respond?

    sessions
  • 58

    A regional Private Service Connect producer onboards 45 consumer projects in ten minutes; its service attachment NAT subnet has four free addresses, 31 new endpoint connections enter PENDING or NEEDS_ATTENTION, and payment API errors reach 18% against a 99.99% SLO; how do you restore consumers?

    slonetworkingprivate-connectivity
  • 59

    Private Cloud DNS SERVFAIL responses jump from 0.005% to 38% across 300 zones after one forwarding change, affecting 50 VPCs with a 99.99% resolution SLO and a five-minute recovery target; what do you do?

    dnsslo
  • 60

    A global external Application Load Balancer returns 22% HTTP 503 at 80,000 requests per second after a backend release, while load-balancer health stays green, connection pools hit their 10,000 limit, and the serving SLO is 99.99%; how do you recover?

    load-balancingslohttp
  • 61

    A regional GKE cluster's Kubernetes API error rate reaches 65% for 20 minutes while existing workloads still serve 99.97% of 70,000 requests per second against a 99.9% SLO; how do you manage the control-plane incident?

    kubernetesapiincidents
  • 62

    A GKE node-pool upgrade applied to two of eight clusters doubles API p99 from 220 ms to 480 ms and raises errors from 0.02% to 3%, with a 99.95% SLO and 5,000 nodes still queued for upgrade; what do you do?

    kubernetesapidata-structures
  • 63

    After a GKE node-image update, packet loss rises from 0.02% to 7% on 900 of 3,600 nodes and service p99 exceeds its 300 ms objective, with 25 minutes of error budget left; how do you isolate a kernel regression?

    kubernetesreliability
  • 64

    A bad external metric drives an HPA from 30 to 1,200 replicas in eight minutes, increases projected compute cost to $40,000 per day, and pushes a database from 45% to 95% utilization against a 99.9% API SLO; how do you stop the runaway?

    databasereplicationapi
  • 65

    A GKE upgrade has drained zero of 600 nodes in 45 minutes because 140 workloads have PDBs with minAvailable equal to replicas, while the upgrade deadline is 24 hours and the serving SLO is 99.95%; what do you do?

    kubernetesreplicationslo
  • 66

    Enabling strict mTLS for 120 services causes 35% HTTP 503 at 60,000 requests per second because 18 legacy workloads lack sidecars, breaching a 99.99% SLO within four minutes; how do you recover?

    mtlshttpslo
  • 67

    A Multi-cluster Gateway route change sends 30% of EU checkout traffic to a US cluster, raising p99 from 240 ms to 1.9 seconds at 45,000 requests per second while the SLO is 99.95%; how do you fix the route incident?

    gatewaysloincidents
  • 68

    Binary Authorization blocks an emergency image during an active exploit, 12 critical services must deploy within a 20-minute mitigation SLO, and the normal attestation pipeline takes 45 minutes; how do you use break-glass safely?

    authsupply-chaindeployment
  • 69

    A compromised Artifact Registry image digest is running in 17 services across six GKE clusters, suspicious egress is 2% above a 20 Gbps baseline, and containment must finish in 15 minutes; how do you respond?

    artifactsregistrieskubernetes
  • 70

    After a shared client rollout, 35 Cloud Run callers mint service-to-service ID tokens with aud set to api.example.com while 12 private targets accept only their run.app URL; 68% of 20,000 requests per second return 401 against a 99.95% SLO, so how do you recover?

    sloserverless-containersapi
  • 71

    A planned AlloyDB primary switchover completes in 70 seconds, but shared connection DNS sends writer traffic to a read pool; 38% of writes fail, read p99 reaches 2.1 seconds, and the application handles 28,000 transactions per second with a five-minute RTO, so what do you do?

    transactionsdns
  • 72

    After an IAM cleanup in a destination project, 96% of eight million daily events from 18 source projects stop reaching a central private Cloud Run service; Eventarc freshness has a two-minute SLO and recovery has a 15-minute RTO, so how do you restore cross-project delivery?

    sloserverless-containersevents
  • 73

    A bad Pub/Sub subscriber release acknowledges messages before committing writes for 12 minutes at 80,000 messages per second, leaving 57.6 million missing records; a snapshot from five minutes before rollout exists, retention is seven days, RTO is 30 minutes, and duplicate effects must remain zero, so how do you recover?

    retentionsnapshotmessaging
  • 74

    One malformed Pub/Sub message triggers a retry storm that consumes 35% of subscriber CPU for 12 minutes, lifts redelivery from 1% to 48%, and threatens a five-minute processing SLO for 4 million valid messages; what do you do?

    resilienceslomessaging
  • 75

    One tenant publishes 1,200 Pub/Sub messages per second averaging 1 KB under one ordering key, exceeding the ordering-key throughput of about 1 MB/s; its lag reaches 35 minutes while product requires five-minute freshness and total tenant ordering, so what do you challenge and change?

    messagingthroughput
  • 76

    A Dataflow release groups one million 1 KB events per second by customer_id; one tenant owns 38%, worker OOM restarts reach 900 per hour, shuffle backlog reaches 6 TB, and p99 freshness rises from 30 seconds to 15 minutes, so how do you recover?

    backlog
  • 77

    On-demand BigQuery scanned bytes rise from 200 TB to 1.6 PB per day after a dashboard release, projected monthly analysis cost grows from $40,000 to about $300,000, and analysts need service restored within 30 minutes; how do you contain the query-cost explosion?

    queriesbigquery
  • 78

    A governance rollout attaches policy tags to 420 BigQuery columns and replaces row access policies on 180 tables; 52 of 60 teams receive Access Denied, eight see zero rows, and 240 dashboards fail with a 30-minute recovery SLO, so how do you recover without overexposing data?

    schemabigqueryslo
  • 79

    A Cloud SQL regional HA failover takes 11 minutes against a two-minute RTO while the database handles 18,000 transactions per second, application errors reach 32%, and the normal baseline is 0.1%; how do you diagnose and fix it?

    sqldatabasetransactions
  • 80

    A 10 TB Cloud SQL for PostgreSQL instance grows from 8.2 TB to 9.9 TB during a runaway load, automatic storage increase is enabled but capped at 10 TB, free space falls to 0.8%, and 12% of writes fail at 22,000 transactions per second; how do you recover?

    sqltransactionspostgres
  • 81

    At 14:03 an operator drops a 1.8 TB orders database on Cloud SQL for PostgreSQL; detection takes eight minutes, accepted-order RPO is zero, PITR is enabled, traffic is 22,000 writes per second, and RTO is 45 minutes, so how do you recover and cut over?

    sqldatabasepostgres
  • 82

    After a settlement batch starts, Spanner ABORTED responses rise from 0.4% to 31% and commit p99 from 90 ms to 1.4 seconds at 45,000 transactions per second; transaction statistics show 12 account rows contended and batch transactions hold locks for four seconds, so how do you recover?

    transactionsbatch
  • 83

    A Spanner schema rollout drops a compatibility column while 30% of 600 clients still read it, causing 9% request failures from a 0.1% baseline at 40,000 requests per second; the SLO allows five minutes to rollback, so what do you do?

    schemarollbackslo
  • 84

    A Cloud Storage lifecycle change reduces object age from 365 to 30 days and deletes 18 TB of required audit exports in two hours, while policy requires seven-year retention and restore RTO is four hours; how do you recover?

    retentionobject-storage
  • 85

    A 90-day CMEK rotation automation switches 80 resources to a replacement CryptoKey whose service-agent grants are missing, so 62% of new writes fail while old reads still work against a 99.99% SLO; how is this different from a disabled-key outage?

    slo
  • 86

    A GCP region loses 85% of capacity for a commerce platform serving 70,000 requests per second, errors reach 28%, the SLO is 99.99%, RTO is 15 minutes, and RPO is five minutes; how do you activate disaster recovery?

    capacityslo
  • 87

    A quarterly restore test for 40 TB of production data finds that the last 14 daily backups all fail checksum validation or cannot access their KMS keys, despite a four-hour RTO and a 15-minute RPO; what do you do?

    validationbackups
  • 88

    Three parallel Terraform applies bypass GCS backend protection because one job uses -lock=false, a second backend prefix manages the same resources, and a break-glass account manually uploads state; the canonical state is truncated and the next plan proposes 600 creates, so how do you recover within 30 minutes?

    terraform
  • 89

    An IAM cleanup removes roles/cloudkms.cryptoKeyEncrypterDecrypter from the Cloud Storage service agent on the CMEK protecting 190 Terraform state objects; 48 teams receive KMS-related 403 errors and the plan service has a 15-minute SLO, so how do you restore access safely?

    terraformsloobject-storage
  • 90

    A semantically invalid Policy Controller constraint reaches 15 GKE clusters through Config Sync and rejects 91% of deployments within six minutes; 3,200 objects still show SYNCED, rollout recovery SLO is ten minutes, and the bad policy also matches platform namespaces, so how do you roll back safely?

    deploymentconfigrollback
  • 91

    A compromised Cloud Build service account publishes backdoored images to nine Artifact Registry repositories and 23 deployments consume them within 25 minutes, while supply-chain containment has a 15-minute SLO; how do you respond?

    deploymentartifactsregistries
  • 92

    A fleet NetworkPolicy rollout blocks Managed Service for Prometheus collectors in 18 of 24 GKE clusters from scraping 9,000 targets and writing samples; ingestion falls 82% for 12 minutes, SLO paging is blind, and observability RTO is ten minutes, so how do you recover?

    observabilitykubernetesmonitoring
  • 93

    Application logs expose email addresses and bearer tokens in six million Cloud Logging entries over 48 hours and 12 sinks export them, while policy requires containment in 30 minutes and confirmed deletion within seven days; how do you respond?

    tokenslogging
  • 94

    One dependency failure triggers 600 pages in 20 minutes from 80 Cloud Monitoring policies, median acknowledgement grows from two to 14 minutes, and the on-call SLO is five minutes; how do you control the alert storm?

    monitoringdependenciesalerting
  • 95

    A 30% canary consumes 22% of a monthly 99.95% error budget in 40 minutes, baseline burn rate was 0.4 and the release gate allows at most 2% budget consumption in one hour; what do you do?

    reliabilitydeployment-strategies
  • 96

    Raw detailed billing export and the invoice show $2.0 million for the month, but chargeback reports only $1.6 million because an effective-dated owner join drops 18% of rows across 320 projects; $400,000 is hidden from team statements, books close in 48 hours, and attribution must exceed 99%, so how do you recover?

    joins
  • 97

    A three-year committed use discount costs $900,000 per year, eligible usage falls 45% after a platform migration, and finance forecasts $250,000 annual waste while serving SLO and traffic remain unchanged; how do you handle the overcommit?

    soft-skillsmigrationsslo
  • 98

    A launch in 72 hours must raise traffic from 20,000 to 200,000 requests per second, but a regional API quota is approved for only 50,000 and the launch SLO is 99.9% with a $60,000 burst budget; what do you do?

    sloapi
  • 99

    During a 25 TB database migration at 40,000 transactions per second, 30% cutover traffic shows 0.8% row mismatches against a 0.1% rollback gate, RPO is zero, and only 20 minutes remain in the window; how do you roll back?

    databasetransactionsmigrations
  • 100

    A major GCP regional incident drives 35% errors across four products serving 120,000 requests per second, with a 99.99% SLO, 20-minute RTO, and five-minute RPO; run incident command with exact communication cadence and recovery gates.

    incidentsslo