Skip to content

Cloud Architect interview questions

100 real questions with model answers and explanations for Senior candidates.

See a Cloud Architect resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

designlanding-zone

I would use a shallow OU hierarchy that separates lifecycle and control requirements, not mirror the company org chart.

  • Put Security, Infrastructure, Workloads, Sandbox, and Suspended directly under the root, then create Production and NonProduction OUs under Workloads.
  • Give each acquisition its own Quarantine OU with SCPs denying organization exit, disabling CloudTrail, and use of unapproved Regions while integration is pending.
  • Keep the root SCP minimal because an inherited deny affects all 300 accounts, and attach narrower guardrails to the lowest OU that needs them.
  • Test every SCP against a canary account with IAM Access Analyzer and service last-accessed data before moving accounts in batches of 10.

Why interviewers ask this: The interviewer is checking whether the candidate understands OU inheritance, SCP blast radius, and safe organization changes at scale.

designcloud-regions

I would separate geography from workload environment so residency rules and production controls can be inherited independently without duplicating every assignment.

  • Under the tenant root, create Platform, Landing Zones, Sandbox, and Decommissioned management groups, with EU and NonEU below Landing Zones.
  • Place Corp-Prod, Corp-NonProd, Online-Prod, and Online-NonProd below each geography, then move subscriptions according to data boundary and environment.
  • Assign an allowed-locations Azure Policy initiative at EU with only approved EU regions, and assign stronger diagnostic and SKU policies at the Prod groups.
  • Use Policy exemptions with an owner and expiry date rather than cloning initiatives, and validate changes in a dedicated canary subscription before 180-subscription rollout.

Why interviewers ask this: The interviewer is evaluating practical Azure hierarchy design, Policy inheritance, and controlled exceptions under residency constraints.

css

I would use folders for inherited controls and projects as the workload, quota, and billing isolation boundary.

  • Create top-level Platform, Products, Sandbox, and Retired folders, then give each product team Production and NonProduction child folders rather than one folder per application.
  • Use separate projects for independently deployed services or environments that need distinct quotas, service accounts, or chargeback labels, keeping shared resources out of product projects.
  • Attach Organization Policy constraints at the highest safe folder and apply labels such as cost-center, product, environment, and owner through the project factory.
  • Link projects to centrally managed billing accounts through automation and export Cloud Billing data to BigQuery for team-level allocation across all 2,000 projects.

Why interviewers ask this: The interviewer is checking whether the candidate can use GCP hierarchy and project boundaries for inheritance, quotas, and cost allocation at scale.

I would make account creation an asynchronous product backed by AWS Control Tower Account Factory for Terraform and a versioned request schema.

  • A pull request supplies owner, cost center, environment, data classification, and target OU, with JSON Schema checks rejecting incomplete requests before merge.
  • Account Factory creates the account, registers it with Control Tower, and applies baseline Terraform for IAM Identity Center, budgets, Config aggregation, and required tags.
  • A second stage attaches the account VPC to the correct AWS Transit Gateway route table through AWS RAM and registers private DNS zones with Route 53 Profiles.
  • The workflow returns the account ID and pipeline status, retries idempotently by request ID, and alerts the platform queue if the 30-minute SLO is exceeded.

Why interviewers ask this: The interviewer is evaluating whether account vending is automated, idempotent, observable, and able to include network and governance baselines.

networkingconcurrency

I would expose a validated subscription request contract and let one pipeline own subscription creation, placement, and connectivity.

  • The request captures workload name, environment, management group, billing scope, owner group, region, and expected address size, with allowed values sourced from a central catalog.
  • The pipeline creates the subscription through an Azure subscription alias, moves it to the target management group, and waits for inherited Policy assignments before deployment continues.
  • An IPAM step reserves a non-overlapping CIDR from Azure Virtual Network Manager IP address management and deploys a spoke VNet peered to the regional hub.
  • Terraform records the subscription ID and allocation, while failed runs can resume from the existing alias instead of producing duplicate subscriptions.

Why interviewers ask this: The interviewer is checking for a concrete, repeatable Azure subscription factory that prevents billing, hierarchy, and address-allocation errors.

networking

I would keep networking in regional Shared VPC host projects and vend analytics projects only as service projects.

  • A Cloud Build pipeline creates each project, assigns its folder and billing account, enables the approved APIs, and waits for project creation propagation before later steps.
  • The factory attaches the project to one of four regional host projects and grants Network User only on the assigned subnets, not at the host-project level.
  • Product owners receive roles inside their service project, while the network team alone manages subnets, routes, Cloud NAT, and firewall policies in the host projects.
  • Project metadata includes owner, cost center, region, and expiry, and the same request key makes retries reconcile the existing project rather than create another one.

Why interviewers ask this: The interviewer is evaluating GCP project automation and the separation of service-project ownership from Shared VPC administration.

dnscloud-regionslanding-zone

I would centralize services that require a single network view while keeping deployment-critical tooling inside each workload account.

  • Put Transit Gateway, Direct Connect attachments, and egress inspection in dedicated Network accounts per Region, with route-table separation for production and nonproduction.
  • Run Route 53 Resolver endpoints and shared private namespace associations in DNS accounts, distributing rules and Route 53 Profiles through AWS RAM.
  • Keep application pipelines, artifact pull-through caches, and runtime secrets account-local so loss of a shared account does not stop a rollback or routine deployment.
  • Use separate Log Archive and Network accounts because combining them would couple high-volume log access with privileged routing changes across 250 accounts.

Why interviewers ask this: The interviewer is checking whether shared services are centralized only where needed and split to limit dependencies and administrative blast radius.

terraformconsistencypartitioning

I would partition state by control plane and account or subscription so one plan never has write authority over all 600 environments.

  • Keep organization-wide hierarchy and policy assignments in a small foundation state changed only by the platform team.
  • Give each account or subscription baseline its own state key, and split regional network hubs into separate states because their lifecycle and operators differ.
  • Store state in a versioned backend with locking, such as S3 plus DynamoDB locking or Azure Storage blob leases, and give each pipeline identity access only to its state path.
  • Pass account IDs, network IDs, and policy versions through a published artifact or parameter registry instead of broad terraform_remote_state access.

Why interviewers ask this: The interviewer is evaluating Terraform state boundaries, concurrency, least blast radius, and explicit dependencies at enterprise scale.

terraformci-cddesign

I would separate reusable module validation from environment plans and serialize applies only at each production state boundary.

  • Pull requests run terraform fmt, validate, TFLint, Checkov, and module tests, then generate a saved plan using the exact provider lock file.
  • Nonproduction account states use independent runners and locks, so up to 10 plans can execute concurrently without sharing credentials or state.
  • Production requires approval of the saved plan, uses an environment-scoped identity with short-lived OIDC credentials, and enforces one apply per state through the backend lock and CI concurrency group.
  • Module releases are pinned by immutable Git tag, and promotion changes only the version reference rather than rebuilding different code for production.

Why interviewers ask this: The interviewer is checking whether the candidate can combine safe Terraform promotion with useful parallelism and deterministic production applies.

linux

I would treat each account move as a policy change and prove effective permissions before moving production accounts.

  • Export current OU membership and attached SCPs, then calculate inherited policies for every account instead of assuming a direct attachment is the full picture.
  • Build target OUs with no new denies first, move two sandbox accounts, and use CloudTrail plus IAM Access Analyzer policy checks to compare expected service calls.
  • Add target SCPs in audit-oriented stages where possible, then move nonproduction accounts in batches of 10 and production accounts by workload maintenance window.
  • Keep a scripted rollback that moves an account to its original parent, and stop the wave if denied API calls or Control Tower drift appears.

Why interviewers ask this: The interviewer is evaluating whether the candidate understands that OU moves change inherited SCPs and needs a measurable, reversible rollout.

designcloud-regionsmulti-region

I would run the stateless checkout tier active-active but keep each inventory item under a single regional writer to preserve correctness.

  • Route users by latency through AWS Global Accelerator to EKS clusters in us-east-1 and eu-west-1, with each checkout tier sized for the full 40,000 requests per second so a regional loss does not turn recovery into a capacity race.
  • Store carts in DynamoDB global tables because sub-second eventual convergence is acceptable, but assign inventory partitions to one home Region and use strongly consistent conditional writes there to prevent two successful reservations.
  • Replicate order events with Amazon MSK Replicator and persist orders in Aurora Global Database; typical cross-Region replication is under a second, but the 30-second RPO must be measured and enforced as a promotion gate.
  • Active-active compute roughly doubles the steady EKS and NAT spend, while single-writer inventory adds about 70 to 100 ms to cross-Atlantic checkout calls; I accept that latency for stock correctness.

Why interviewers ask this: The interviewer is testing whether the candidate separates workloads by consistency need instead of claiming that active-active makes every component both writable and strongly consistent.

cloud-regionscloud-modelscloud

I would pin every tenant to a residency-scoped cell and replicate only non-customer control metadata globally.

  • Deploy identical Azure landing-zone cells in West Europe, North Europe, Canada Central, and Canada East, each with AKS, Azure SQL failover groups, and a tenant registry that maps the tenant ID to its legal geography.
  • Keep customer records, backups, logs, and encryption keys inside the paired geography using Azure Policy and region-specific Key Vaults; Azure Front Door may route traffic globally, but it must not cache regulated response bodies outside that geography.
  • Size each secondary cell for 35% of regional peak and use autoscaling to reach 100% within 20 minutes, which meets the 30-minute RTO at materially lower cost than fully active duplicate capacity.
  • Configure SQL failover groups and geo-redundant storage for a tested RPO below 5 minutes, accepting asynchronous replication and up to 5 minutes of acknowledged-write loss rather than violating residency with a global database.

Why interviewers ask this: A strong answer turns residency into a routing, storage, key, log, and backup boundary while quantifying the capacity needed for recovery.

system-designdesigncloud-regions

I would use a 25% warm standby with continuously replicated data and pre-provisioned network, identity, and deployment infrastructure.

  • Run the primary on ECS Fargate and Aurora PostgreSQL in us-east-1, with a scaled-down ECS service and an Aurora Global Database secondary in us-west-2; Application Auto Scaling must restore full task capacity within 30 minutes.
  • Replicate documents with S3 Cross-Region Replication and use S3 Replication Time Control for the critical bucket, whose 15-minute replication SLA matches the RPO but adds replication and per-GB transfer charges.
  • Pre-create VPC endpoints, KMS multi-Region keys, IAM roles, Secrets Manager replicas, and Route 53 records so the 60-minute RTO does not depend on creating infrastructure during failover.
  • Keeping compute at 25% plus the database secondary should fit the 35% cap, but Aurora, S3 RTC, and inter-Region transfer remain fixed costs; I would reject pilot light because image pulls and database provisioning make the 60-minute RTO fragile.

Why interviewers ask this: The scenario checks whether the candidate can translate a cost ceiling and recovery objectives into a specific warm-standby capacity level.

designdnsfailover

I would put an anycast global proxy in front of both regions rather than make DNS TTL the primary failover mechanism.

  • Use Azure Front Door Premium with health probes every 30 seconds and three failed probes before removing an origin, giving roughly 90 seconds for detection while keeping the hostname and edge IP stable for clients.
  • Run active-active Application Gateway and AKS stacks in both regions, with each region able to serve at least 60% of 15,000 requests per second and Front Door weighted 50:50 during normal operation.
  • Keep Azure Traffic Manager only as a fallback for the Front Door endpoint, because recursive resolvers and the stated 10-minute client cache make DNS-only failover unable to meet a 3-minute RTO.
  • Front Door Premium adds a global proxy and data-processing bill, and cross-region overflow increases egress and user latency by about 80 to 120 ms, but it buys failover independent of client DNS behavior.

Why interviewers ask this: The interviewer wants the candidate to recognize that a low DNS TTL is not an enforceable recovery mechanism when clients cache longer.

designserializationcloud-regions

I would use a multi-region Spanner database and make the application tolerate its quorum latency instead of weakening ledger consistency.

  • Place Cloud Run or GKE services in us-central1 and europe-west4 behind a global external Application Load Balancer, then use Spanner multi-region configuration nam-eur-asia1 or a custom instance configuration aligned to the two traffic centers.
  • Use Spanner read-write transactions and commit timestamps for balance changes; synchronous Paxos replication gives zero RPO for committed transactions but cross-continent quorum directly consumes the 180 ms p99 budget.
  • Send balance reads through strong reads and route noncritical history queries to bounded-staleness reads, such as 15 seconds, so reporting does not compete with the 8,000 writes per second.
  • Multi-region Spanner requires more replica compute than a regional instance and charges inter-region network traffic; if load tests exceed 180 ms, I would move the voting topology closer to the dominant writer rather than introduce dual writable ledgers.

Why interviewers ask this: The question tests whether the candidate protects financial invariants and treats quorum placement and latency as measurable architecture constraints.

designbackupsvalidation

I would keep immutable backups in a second EU Region and size the restore path from the 40 TB in 24 hours requirement.

  • Use AWS Backup for application-consistent snapshots, copy them from eu-central-1 to eu-west-1 every 4 hours, and lock the destination vault with AWS Backup Vault Lock so retention cannot be shortened.
  • Store archive objects with S3 Versioning and EU-only Cross-Region Replication, then transition older versions to S3 Glacier Deep Archive; the low storage price trades off against up to 12 hours for standard retrieval.
  • Restoring 40 TB in 24 hours requires sustained throughput of about 470 MB/s before overhead, so I would provision parallel S3 reads, sufficient EC2 and EBS bandwidth, and a 30% margin rather than assume storage retrieval is the only bottleneck.
  • Run a quarterly full restore into an isolated eu-west-1 account and measure database recovery plus checksum validation; duplicate storage, retrieval requests, and temporary compute are explicit compliance costs.

Why interviewers ask this: A strong answer derives restore throughput from the deadline and includes residency, immutability, and a full restoration test rather than merely naming a backup service.

dependencies

I would split each region into independently deployable 10% tenant cells and give every tenant a deterministic home cell.

  • Build ten cells across us-east-1, us-west-2, and eu-west-1, each with its own EKS node groups, Aurora cluster, SQS queues, and per-cell quotas sized for about 6,000 requests per second.
  • Keep a small DynamoDB global table as the tenant-to-cell directory and cache it at CloudFront; the directory contains routing metadata only, so a cell cannot access another cell's customer database.
  • Replicate each cell asynchronously to a paired standby cell with Aurora Global Database and S3 Cross-Region Replication, alarming when lag approaches 60 seconds and promoting within the 10-minute RTO.
  • Ten database and queue stacks cost more and leave roughly 15% spare capacity stranded per cell, but they cap blast radius and let one cell move without scaling or failing over all 30 million users.

Why interviewers ask this: The interviewer is evaluating whether the candidate can make a numerical blast-radius goal concrete through isolation, routing, capacity, and replication boundaries.

design

I would separate residency-bound source objects from globally distributable derivatives at the bucket and pipeline boundaries.

  • Send creators through a global external Application Load Balancer to regional upload services, then write EU source files to an EU dual-region Cloud Storage bucket with public access prevention and CMEK keys kept in EU Cloud KMS.
  • Run Transcoder API or GKE workers in EU for EU sources, write only approved derivatives to a separate multi-region Cloud Storage bucket, and serve those through Cloud CDN so the 6 PB monthly delivery does not return to origins.
  • Dual-region storage provides regional redundancy without exporting raw files outside the EU; asynchronous replication means I would buffer upload metadata in regional Pub/Sub and set the source workflow RPO to the measured replication lag, capped at 5 minutes.
  • Keeping source and derivative copies increases storage, operation, and replication charges, while global CDN egress dominates delivery cost; lifecycle raw intermediates after 30 days to control the duplicate footprint.

Why interviewers ask this: The question checks whether the candidate enforces residency on the raw-data path while allowing a separately governed public artifact to use global delivery.

databasecloud-regions

I would not make Aurora active-active for seat writes; I would use one writer Region with active-active read and compute tiers.

  • Route browsing to local EKS services and Aurora readers through AWS Global Accelerator, but send seat reservations to the current Aurora PostgreSQL writer and protect them with a unique seat constraint plus a serializable transaction.
  • Use Aurora Global Database managed planned failover or switchover, monitor GlobalDBReplicationLag below 1 second, and block promotion if the lag breaches the stated RPO rather than silently losing accepted reservations.
  • Pre-scale the secondary EKS cluster and database instance to carry all 3,000 writes per second so DNS and connection-pool changes can complete inside the 5-minute RTO.
  • Cross-Atlantic writes add roughly 70 to 100 ms for users away from the writer and duplicate database capacity raises cost, but application-level multi-writer conflict resolution is unsafe for exclusive seats.

Why interviewers ask this: The interviewer is looking for a clear rejection of multi-writer semantics where a unique business resource requires one serialization point.

backupsfailoverdisaster-recovery

I would combine a small warm application tier with continuously replicated state, while treating immutable backups as the fallback for corruption rather than the primary two-hour recovery path.

  • Run the primary in Australia East and a 15% AKS warm pool in Australia Southeast, with container images in geo-replicated Azure Container Registry and autoscaling tested to reach batch capacity within 60 minutes.
  • Use Azure SQL failover groups with a 15-minute data-loss ceiling and geo-redundant Storage restricted to Australian regions; keep payroll keys in separate regional Key Vaults under the same residency boundary.
  • Store daily immutable Azure Backup recovery points for 35 days and perform a monthly restore of all 4 million records, but rely on SQL replication for the two-hour RTO because a full database restore may consume most of that window.
  • A 15% warm pool plus secondary SQL compute must fit the 20% premium, so I would pause nonessential reporting during failover; fully active compute would improve RTO but exceed the cost cap for a six-hour workload.

Why interviewers ask this: This scenario tests whether the candidate distinguishes rapid regional failover from slower backup restoration while honoring residency and a strict cost ceiling.

Locked questions

  • 21

    You are designing AWS networking for 12 new VPCs across three Availability Zones, with up to 2,000 instances per VPC and a future on-premises connection using 10.0.0.0/8. How would you allocate addresses and routes without creating an early scaling limit?

    scalingdesignnetworking
  • 22

    After an acquisition, 18 AWS VPCs and two data centers all use parts of 10.20.0.0/16, but 40 applications must communicate within six months and renumbering them will take two years. What networking pattern would you use?

    communication
  • 23

    Forty AWS VPCs connect to one Transit Gateway: 12 production VPCs may reach shared DNS and logging, 20 nonproduction VPCs must not reach production, and eight shared-service VPCs need selected return paths. How would you build the route segmentation?

    gatewaydnslogging
  • 24

    A payments API in one AWS VPC must be consumed privately by 60 VPCs owned by different business units, peaks at 2 Gbps, and must not expose any other provider subnets. Would you choose PrivateLink, VPC peering, or NAT?

    networkingapi
  • 25

    A data center must send 6 Gbps continuously to 24 AWS VPCs through Transit Gateway with under 25 ms network latency, and the business allows no single carrier or edge-location failure. How would you design Direct Connect and VPN routing?

    designgatewaylatency
  • 26

    A company operates 25 Azure spokes and 25 GCP VPCs, needs centralized firewalls, 5 Gbps private connectivity to two data centers, and isolation between production and development. What equivalents would you use to implement the same hub-and-spoke intent in both clouds?

    networking
  • 27

    Eighty AWS VPCs and two data centers must resolve private names in corp.example and service.aws.internal at a peak of 25,000 DNS queries per second, with no public leakage or forwarding loops. How would you design DNS?

    designdnsqueries
  • 28

    An AWS application transfers 8 TB per day between app and database tiers across Availability Zones at an illustrative $0.01 per GB per charged direction, plus 3 TB per day to a region 70 ms away at $0.02 per GB. How would you reduce transfer cost without weakening the recovery target?

    availabilitycloud-regionshigh-availability
  • 29

    Twenty AWS VPCs send 40 Gbps of patch and container-image traffic to the internet, and centralized NAT gateways are producing high processing and cross-AZ charges. How would you redesign egress for capacity and cost?

    gatewaycapacitynetworking
  • 30

    Thirty AWS VPCs must send 15 Gbps of east-west and internet traffic through stateful third-party firewalls, but asymmetric return traffic currently causes sessions to drop. How would you design the inspection path?

    designnetworkingsessions
  • 31

    A B2B ordering service needs relational constraints and multi-row transactions, serves 6,000 writes and 25,000 reads per second, holds 8 TB, grows by 2 TB yearly, and requires p99 below 80 ms, RPO under 1 minute, and RTO under 15 minutes. Which database would you choose?

    databasetransactions
  • 32

    An IoT platform receives 300,000 device readings per second, each under 2 KB, retains 30 days online, reads almost exclusively by device ID and timestamp, and can tolerate 200 ms p99 with no joins or cross-device transactions. What storage would you select?

    transactionsjoins
  • 33

    A global reservation system runs in North America, Europe, and Asia, performs 40,000 transactions per second across inventory and booking rows, must prevent double booking, and needs local read p99 under 120 ms with a regional RTO below 5 minutes. Would you choose Cloud Spanner or Azure Cosmos DB?

    system-designcloud-regionstransactions
  • 34

    A media company stores 6 PB of immutable video masters, ingests 10 TB daily, retrieves 2 percent of objects each month, must begin restoring any archived title within 4 hours, and requires protection from accidental deletion for seven years. How would you store it?

    immutability
  • 35

    A product analytics platform ingests 20 TB of events per day, retains 3 PB for seven years, needs executive dashboards to return in under 10 seconds, and lets data scientists scan arbitrary history a few times per week. Would you build a warehouse or a lakehouse?

    warehouse
  • 36

    Two hundred rendering nodes must share a 500 TB POSIX namespace, read the same large scene files concurrently at up to 80 GB/s, write temporary outputs, and release capacity between nightly jobs. Would you use block or file storage?

    capacitycloud-storageconcurrency
  • 37

    A self-managed PostgreSQL database on EC2 is 12 TB and generates 70,000 random 16 KB IOPS at peak, needs storage p99 below 2 ms, sustains about 1.1 GB/s, and grows 20 percent yearly. Which EBS volume design would you choose?

    databasepostgresdesign
  • 38

    A 40 TB Aurora PostgreSQL payments database requires RPO below 5 minutes and RTO below 30 minutes even after complete regional loss, while backups must be immutable for 35 days. What recovery storage design would you implement?

    designbackupscloud-regions
  • 39

    A shopping cart service serves 120,000 reads and 20,000 writes per second across six Azure regions, requires users to read their own writes within 100 ms, tolerates other users seeing data up to 2 seconds late, and must remain writable during a regional partition. Which Cosmos DB consistency level would you use?

    consistencycloud-regionspartitioning
  • 40

    A compliance platform has a 70 TB PostgreSQL audit table growing by 2,000 append-only rows per second, reads the latest 90 days at 500 queries per second, runs fewer than 20 yearly investigations against older data, and must retain records for seven years. Would you keep all data in PostgreSQL?

    postgresqueries
  • 41

    Your AWS compute bill is $420,000 per month, 70% of usage has been stable for a year, and product demand can still swing by 25%; how much would you commit to with Savings Plans?

  • 42

    Nonproduction EC2, RDS, and EKS resources cost $96,000 per month, but developers use them only from 08:00 to 20:00 on weekdays; what would you change without blocking testing?

    kubernetesawstesting
  • 43

    An EKS platform spends $180,000 per month on worker nodes; 65% of CPU runs stateless jobs, Spot is 60% cheaper, and the service must retain 99.95% availability during interruptions, so how would you redesign capacity?

    capacitykubernetes
  • 44

    A new image processor runs 20 million jobs per month, each using 4 GB for 10 seconds; Lambda costs $0.000016667 per GB-second, while 60 c7g.xlarge ECS instances cost $0.145 per hour, so which platform do you choose?

    concurrencylambdacompute
  • 45

    Eight petabytes of audit logs sit in S3 Standard at about $188,000 per month; only 15% is read after 30 days, retention is one year, and retrieval must finish within hours, so what lifecycle policy would you apply?

    retentionaws
  • 46

    A media service moves 900 TB per month across Availability Zones and 300 TB to users, producing a $45,000 transfer bill; which architecture changes would you fund first?

    availabilityhigh-availability
  • 47

    Datadog costs $140,000 per month for 60 TB of logs and full tracing, but security requires 90-day retention for authentication events and engineers need seven days of searchable application logs; how would you cut the bill?

    retentionauthmonitoring
  • 48

    A commerce platform spends $310,000 per month in cloud costs for 2.5 million completed orders, and finance requires infrastructure cost below $0.10 per order before a low-margin market launch; what do you do?

    css
  • 49

    A single-region SaaS stack costs $240,000 per month, requires regional recovery within four hours, and averages six hours of regional downtime per year at $40,000 per hour; active-active would add 85% while warm standby adds 35% and cuts downtime to two hours, so what do you choose?

    cloud-regionscloud-modelscloud
  • 50

    A payments database costs $75,000 per month in one AWS region; the contract requires RPO under 5 minutes and RTO under 30 minutes, backups restore in four hours, a warm cross-region Aurora cluster adds $12,000, and active-active adds $38,000, so what do you build?

    backupscloud-regionsdatabase
  • 51

    An AWS checkout platform handles 22,000 requests per second on EKS and Aurora Global Database across us-east-1 and us-west-2, with RTO 5 minutes and RPO 30 seconds; us-east-1 failed 11 minutes ago, promotion finished at minute 4, but Global Accelerator still sends 35% of traffic to unhealthy pods and recovery has exceeded RTO. What do you do now, and how do you restore service safely?

    databasekubernetes
  • 52

    An Azure order service processes 6,000 writes per second with AKS and Azure SQL Managed Instance link between East US 2 and Central US, with RTO 10 minutes and RPO 1 minute; after a network partition both regions accepted writes for 7 minutes, producing 84,000 conflicting orders even though observed recovery was 8 minutes. What do you do now, and how do you restore from this split-brain safely?

    sqlpartitioningcloud-regions
  • 53

    A GCP logistics API serves 14,000 requests per second on GKE with a 9 TB Cloud SQL for PostgreSQL database replicated from europe-west1 to europe-west4, with RTO 15 minutes and RPO 2 minutes; the primary is unavailable, the replica is 18 minutes behind, and the incident has lasted 6 minutes. What do you do now, and how do you restore safely despite the replication lag?

    sqldatabasepostgres
  • 54

    An AWS SaaS API serves 30,000 requests per second on EKS and Aurora Global Database in eu-west-1 and eu-central-1 with RTO 4 minutes; compute and database failover finished in 3 minutes, but 62% of logins still fail because the external identity provider callback and partner allowlist reference the old region's ALB and NAT IPs. What do you do now?

    networkingcloud-regionscloud-models
  • 55

    A multi-region Azure SaaS platform handles 18,000 requests per second with Front Door, AKS, Cosmos DB, and regional Key Vaults, with RTO 8 minutes and RPO 1 minute; West Europe failed 13 minutes ago, North Europe is healthy, but both regions depend on a single private DNS zone link and Azure Firewall in West Europe, so login and secret resolution still fail globally and the 8-minute RTO has been exceeded by 5 minutes. What do you do now, and how do you restore safely?

    dnsnetworkingcloud-regions
  • 56

    An AWS claims system stores a 24 TB Aurora PostgreSQL cluster and 110 TB in S3, with RTO 6 hours and RPO 15 minutes; ransomware requires restoration into a clean account, but after 4 hours AWS Backup has restored only 6 TB of the database and the projected completion is 14 hours, while the chosen recovery point is 11 minutes old and a protected native Aurora snapshot from 13 minutes before the attack is also available. What do you do now, and how do you restore safely?

    databasepostgressystem-design
  • 57

    A GCP commerce platform normally serves 40,000 requests per second on GKE in us-central1 backed by multi-region Spanner, with a 25% warm standby in us-east1, RTO 12 minutes and RPO 0; after regional evacuation, us-east1 reached CPU saturation at 18,000 requests per second, checkout errors are 22%, and observed recovery is 17 minutes. What do you do now, and how do you restore safely with the capacity shortfall?

    capacitycloud-regionsmulti-region
  • 58

    An Azure billing pipeline processes 9,000 Service Bus messages per second into a 14 TB Azure SQL Database, with RTO 20 minutes and RPO 5 minutes; after Central US failover, operators replayed a 12-million-message queue, duplicate writes reached 310,000 invoices, and service resumed in 16 minutes with no message loss. What do you do now, and how do you restore correctness safely?

    queuesfailovermessaging
  • 59

    An AWS banking API processes 5,500 writes per second on EKS and Aurora Global Database, with RTO 7 minutes and RPO 30 seconds; eu-west-1 was emergency-promoted 9 hours ago after us-east-1 failed, service recovered in 6 minutes with 12 seconds of lag, and us-east-1 is now available but contains 190 million stale rows. What do you do now, and how do you fail back safely?

    databaseapiconcurrency
  • 60

    A GCP payments service writes 7,500 transactions per second to a 6 TB Cloud SQL for PostgreSQL primary in asia-southeast1 with a cross-region replica in asia-east1, with RTO 10 minutes and RPO 60 seconds; the primary is lost, promotion completed in 8 minutes, but the promoted database is 4 minutes 20 seconds behind and 1.95 million acknowledged transactions may be absent. What do you do now, and how do you restore safely after the RPO breach?

    sqldatabasetransactions
  • 61

    The AWS network budget was $45,000 for the month, but the forecast jumped to $182,000 after a release and finance needs a containment plan in four hours. NAT Gateway charges are the main variance. How do you investigate and cut the bill without breaking production egress?

    gatewaynetworkingdispersion
  • 62

    An EKS platform was planned at $70,000 per month, but it is now forecasting $128,000 after teams adopted Spot nodes and Karpenter. You have one day to find the regression and reduce cost while preserving a 99.95% API target. What do you do?

    decision-makingapikubernetes
  • 63

    Azure Monitor and Log Analytics were budgeted at $25,000 per month, but the forecast is $140,000 after a diagnostic settings rollout. Security gives you six hours to cut ingestion without losing audit evidence. How do you investigate and act?

    monitoring
  • 64

    AWS Compute was expected to cost $330,000 this month, but the bill is forecasting $470,000 and a $300,000 monthly EC2 Instance Savings Plan shows only 52% utilization instead of the planned 92%. Workloads recently moved instance family and Region. What do you investigate and change before tomorrow's CFO review?

    cloud-regionscomputeaws
  • 65

    A GCP serverless pipeline was planned at $12,000 per month, but Cloud Billing now forecasts $96,000 after six hours of repeated failures. Pub/Sub invokes Cloud Run and failed messages retry indefinitely. How do you contain and fix it before the next billing export?

    resilienceserverlessci-cd
  • 66

    After an AWS multi-region launch, API p99 latency rose from the 180 ms objective to 640 ms while p50 stayed near 70 ms. The launch window closes in 90 minutes. How do you investigate and make an immediate go or rollback decision?

    cloud-regionsmulti-regionapi
  • 67

    An Azure Event Hubs ingestion service was load-tested for 120 MB/s, but at the live peak it plateaus near 70 MB/s, returns throttling errors, and the backlog will breach its 15-minute NFR in 25 minutes. How do you investigate and recover?

    backlog
  • 68

    A GCP database migration promised p99 storage latency below 20 ms at 18,000 IOPS, but production on balanced Persistent Disk reaches 240 ms and queue depth keeps rising. You have two hours before the migration must be rolled back. How do you investigate and fix the IOPS shortfall?

    databasemigrationspromises
  • 69

    A GCP customer API committed to 99.95% monthly availability but is at 99.72%, about 121 minutes unavailable in a 30-day month versus a 21.6-minute budget. A regional Cloud Run revision caused most failures and renewal talks start tomorrow. What do you investigate and fix today?

    cloud-regionsapi
  • 70

    An AWS checkout platform costs $520,000 per month against a hard $350,000 cap, and leadership demands a $170,000 cut within 48 hours while keeping the current 99.99% availability claim. The second region is active-active and costs $145,000 monthly. How do you investigate, decide, and execute?

    cloud-regions
  • 71

    During an 11 TB PostgreSQL cutover to Amazon Aurora, AWS DMS CDC lag jumps from 8 seconds to 45 minutes after 20% of orders move to the target. Walk me through your live response, rollback, and safe resumption.

    postgresrollback
  • 72

    A monolith moved from EC2 to ECS shows checkout p99 rising from 180 ms to 950 ms when the Application Load Balancer sends 30% of production traffic to the new tasks. What do you do live, and how do you restart the migration safely?

    monolithload-balancingmicroservices
  • 73

    After 15% of read traffic moves from Azure SQL in East US 2 to a new West Europe deployment, reconciliation finds mismatches in 2.4% of 180 million customer rows. Walk me through containment, rollback, and a safe second attempt.

    deploymentrollbacksql
  • 74

    During a wave of 120 VMware servers into AWS, replication drives a 10 Gbps Direct Connect link to 96% and the on-premises ERP starts timing out. How do you contain the incident, roll back the wave, and resume it?

    replicationincidentsrollback
  • 75

    During migration of 150 batch workers from GCE VMs to GKE, 20% of jobs move to the new workers, but 18% fail with Cloud Storage 403 responses and the Pub/Sub backlog reaches 1.2 million messages against a 15-minute age SLO. Walk me through containment, rollback, and safe resumption.

    batchslocloud-migration
  • 76

    An AWS Config alert shows that a Terraform change made 14,200 private S3 objects publicly readable for 23 minutes in production. What is your immediate response, rollback, and path to safely restore customer downloads?

    terraformconfigrollback
  • 77

    You discover that an EU-only GCP workload exported 180 GB of application logs and three database backups to us-central1 over the last 36 hours. How do you contain the residency breach and resume logging and backups safely?

    databaseloggingbackups
  • 78

    A production AWS KMS key-policy rollout removes the application role and the break-glass admin, causing decrypt failures to jump to 18% across 40 services. Walk me through containment, key-control recovery, rollback, and safe resumption.

    rollback
  • 79

    Forty-eight hours before a SOC 2 audit, you find that AWS Config recording was disabled in 37 production accounts for 19 days after a Control Tower change. What do you do now, how do you roll back the drift, and when do you resume normal releases?

    configrollbackiac
  • 80

    A new Azure Policy initiative with deny effects reaches 42 production subscriptions and blocks AKS node-pool scaling during a traffic peak, leaving 600 pods pending. Walk me through containment, rollback, and a safer policy restart.

    scalingrollback
  • 81

    Your proprietary managed database now costs $162,000 per month after a 35% price rise, up from $120,000, and renewal is due in 60 days. It stores 18 TB, peaks at 7,000 writes per second, and the product can tolerate one 20-minute cutover. What decision do you make now, and how do you execute it?

    database
  • 82

    A proprietary cloud queue is capped at 50,000 messages per second, while measured growth will require 80,000 within four months and the provider has denied a quota increase. The platform carries 25 TB of retained events and permits no producer outage. What do you choose and what do you do this week?

    data-structures
  • 83

    You must move 4 PB from AWS to GCP, but internet egress is $0.05 per GB and the current contract ends in 90 days. The destination can ingest 40 Gbps and the source provider offers no free exit. Do you stay or leave, and how do you execute the decision now?

  • 84

    The provider will discontinue the managed application runtime hosting 70 Java services in six months. The estate serves 12,000 requests per second, has 99.95% availability, and the team can support either Azure Container Apps or AKS but not both. What do you select and how do you exit the retired service?

    containers
  • 85

    A Terraform provider breaking change makes plans propose replacement of 1,800 resources across 420 states, and a security fix in the same provider release must reach production within 14 days. What decision do you make today and how do you roll it out?

    terraformversioning
  • 86

    Six engineers have 12 weeks to launch a regulated order workflow with retries, timers, audit history, and 99.9% availability. Do you build orchestration, adopt AWS Step Functions, use Temporal Cloud, or operate Temporal yourself, and what do you commit to now?

    decision-makingorchestration
  • 87

    A claims portal must launch in 10 weeks at 5,000 requests per second, but 94% of requests are status reads and the current design calls a mainframe policy API capped at 500 calls per second on every request; p99 is 4.2 seconds against an 800 ms SLO. What architecture do you approve now?

    designsloapi
  • 88

    A document processor must launch in six weeks for 200,000 jobs per day; 70% is already implemented on Lambda, but tests show six-minute average jobs, 22-minute p95 duration against Lambda's 15-minute limit, and 4% duplicate outputs after retries. What do you deploy now?

    concurrencylambdadeployment
  • 89

    Security mandates that every payment request pass through a central inspection proxy, but tests show it adds 38 ms to an existing 31 ms p99 and the contractual limit is 50 ms. You have three weeks before launch. What control do you approve and how do you prove it?

    proxytesting
  • 90

    A checkout API must maintain 99.99% monthly availability at 40,000 requests per second, but its recommendation service is failing 30% of calls, loyalty writes are timing out, and inventory is healthy. Which degradation do you activate now, and what do you refuse to degrade?

    api
  • 91

    A junior architect proposes one /16 AWS VPC for 35 services owned by eight teams because the platform must launch in six weeks. What do you say and do, and what concrete artifact or validation do you require before approving the network design?

    designnetworkingvalidation
  • 92

    A senior architect proposes active-active writes across us-east-1 and eu-west-1 for an order service handling 12,000 writes per second, but the design has no conflict model and launch is in eight weeks. What do you say and do, and what artifact or validation must exist before the decision?

    validationdesignartifacts
  • 93

    An architect submits an RFC for a 6,000-request-per-second API with two options: EKS measured at 42 ms p99 and $18,000 per month, and ECS Fargate measured at 55 ms p99 and $24,000 per month; the SLO is 80 ms and a decision is due Friday. What do you say and do, and what concrete evidence do you require?

    decision-makingapislo
  • 94

    Two weeks before launch, your review finds that a 14 TB PostgreSQL database restores in 11 hours although the committed RTO is 60 minutes and the team has no cross-region replica. What do you say and do, and what artifact or validation do you require before launch?

    databasepostgresreplication
  • 95

    An architect has produced three designs for a 50-account AWS platform, each with detailed diagrams but no cost model, and procurement needs a budget in five days. What do you say and do, and what concrete artifact or validation do you require?

    procurementdesignvalidation
  • 96

    Two teams disagree on EKS versus Lambda for a service measured at 900 requests per second normally, 8,000 for ten minutes twice a day, 250 ms average execution, and a 100 ms cold-start budget; launch is in four weeks. What do you say and do, and what artifact or validation do you require?

    artifactskubernetesconflict
  • 97

    A product team requests a 30-day exception to an EU residency control so 4 TB of customer exports can be processed in us-east-1 before a contract deadline in three weeks. What do you say and do, and what artifact or validation do you require?

    concurrencyerror-handlingvalidation
  • 98

    A mentee submits a Terraform change that replaces the IAM baseline module used by 200 AWS accounts, and the release window opens in 48 hours. What do you say and do, and what concrete artifact or validation do you require before apply?

    terraformartifactsvalidation
  • 99

    A postmortem after a 95-minute cloud outage contains 40 actions such as 'improve monitoring' and 'review architecture,' with no owners or dates; the review is tomorrow. What do you say and do, and what concrete artifact or validation do you require?

    monitoringartifactsarchitecture
  • 100

    Only 20% of 60 application teams use the new cloud platform standard because onboarding takes three days, while the target is 70% adoption in one quarter. What do you say and do, and what concrete artifact or validation do you require?

    onboardingdecision-makingvalidation