Skip to content

Solutions Architect interview questions

100 real questions with model answers and explanations for Senior candidates.

See a Solutions Architect resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

designcloud-regionsauth

I would claim each payment key in a globally strong store, then keep the financial write under one regional authority.

  • Amazon Route 53 sends clients to the nearest region, where a conditional write to DynamoDB Global Tables in multi-region strong consistency mode claims the UUID idempotency key before charging.
  • The regional PostgreSQL ledger has the same unique key, the payment provider receives it on every retry, and an outbox publishes PaymentAuthorized only after commit.
  • Each region runs at 60% peak capacity so the other two can absorb a failure; this costs about 50% more compute than active-passive but avoids a 20-minute promotion RTO.

Why interviewers ask this: The interviewer is evaluating whether the design meets quantified latency and correctness goals without pretending global consistency is free.

designcloud-modelscloud

I would use pooled cells for standard tenants and dedicated accounts and databases only for regulated banks.

  • Twenty cells with PostgreSQL row-level security hold about 1,000 standard tenants each, limiting a cell failure to 5% of pooled customers.
  • Each bank receives an AWS account, KMS key, and Aurora cluster, adding roughly $1,200 monthly per bank but satisfying contractual isolation and tenant-level restore.
  • A control plane maps tenant IDs to cells, enforces quotas, and supports online movement through CDC, so a growing tenant can change tier without changing APIs.

Why interviewers ask this: A strong answer balances isolation, failure radius, migration capability, and a concrete budget.

designcloud-regions

I would split regulated source records by jurisdiction and replicate only an approved global read model.

  • EU PII stays in eu-west-1 and US PII in us-east-1, each encrypted with regional AWS KMS keys and exposed through region-bound service endpoints.
  • Kafka topics carry ProfileDisplayNameChanged events without addresses or tax IDs, and regional consumers build Redis read models with a 60-second freshness target.
  • Global replication adds about $9,000 monthly in brokers, cache, and transfer, but replicating full PII would violate residency and expand audit scope.

Why interviewers ask this: The interviewer is checking whether residency controls cover payloads, keys, events, and read models while preserving latency.

oauthapidesign

I would place a versioned REST contract behind Kong Gateway and isolate authentication, quota, and business processing.

  • Kong validates OAuth 2.0 tokens from Okta, applies token-bucket quotas in Redis, and returns RateLimit headers before traffic reaches regional services.
  • OpenAPI is the contract source, additive changes stay in v1, and breaking changes enter v2 with 12 months of dual operation and consumer contract tests.
  • A managed gateway would reduce operations but cost about $90,000 monthly at 25,000 RPS, while Kong on Kubernetes is about $35,000 plus two engineer-months yearly.

Why interviewers ask this: The interviewer is evaluating API lifecycle design, traffic control, and an explicit managed-versus-operated cost trade-off.

designconcurrency

I would commit the order synchronously and move fulfillment to an event-driven workflow.

  • The checkout API writes Order and outbox rows in PostgreSQL within 250 ms, then Debezium publishes OrderPlaced to Kafka without a dual-write gap.
  • Inventory, payment capture, and shipping consumers use order ID as an idempotency key, with Temporal coordinating compensation when a step exhausts retries.
  • Kafka and Temporal add about $28,000 monthly and projection lag, but they keep slow warehouse APIs outside checkout and absorb a 10-times traffic burst.

Why interviewers ask this: A strong answer separates the latency-critical commit from long-running work and addresses atomic publication and compensation.

backlogdesign

I would use separate Amazon SQS queues per service tier with weighted worker pools and explicit admission limits.

  • Premium, standard, and bulk queues receive 60%, 30%, and 10% minimum worker capacity, while idle workers may borrow capacity across tiers.
  • Producers are throttled when bulk age exceeds 30 minutes or its queue reaches 50 million messages, and failed tasks move to named dead-letter queues after five attempts.
  • SQS costs roughly $21,000 monthly at this volume; Kafka could lower per-message cost but adds partition management and weaker native per-message delay handling.

Why interviewers ask this: The interviewer is checking priority isolation, finite backlog policy, failure handling, and queue technology economics.

system-designcloud-migrationmigrations

I would migrate by policy capability behind an anti-corruption API, not replace the mainframe in one cutover.

  • Months 1 to 3 map COBOL rules, calls, and data ownership; months 4 to 8 introduce Apigee and IBM MQ adapters while read-only capabilities move first.
  • Months 9 to 13 use CDC into PostgreSQL, reconcile counts and financial totals daily, then move writes one product line at a time with a four-hour rollback boundary; month 14 removes adapters after 30 stable days.
  • Five months of parallel operation cost about $700,000, but it contains semantic and data-loss risk that a single cutover would concentrate into one outage.

Why interviewers ask this: The interviewer is evaluating staged coexistence, dependency control, measurable reconciliation, and the price of risk reduction.

databasepostgresmigrations

I would combine schema conversion, bulk load, and continuous replication, then transfer write authority during a rehearsed cutover.

  • Months 1 to 2 use AWS SCT to classify incompatible PL/SQL and choose rewrites; months 3 to 5 load snapshots and run AWS DMS CDC with lag below 30 seconds.
  • Months 6 to 8 dual-read shadow traffic, compare row counts, checksums, and 25 financial invariants, then rehearse a 20-minute write freeze and DNS switch twice.
  • Month 9 freezes writes, drains DMS lag to zero, switches authority, and keeps Oracle fenced but recoverable for a four-hour rollback window.
  • Aurora saves an estimated $420,000 yearly in licenses, while the migration costs about $650,000 and risks semantic differences in sequences, isolation, and stored procedures.

Why interviewers ask this: A strong answer includes stages, validation gates, cutover numbers, technology, and full migration economics.

monolithmicroservicesdatabase

I would use a strangler architecture around business slices while the monolith remains the initial system of record.

  • An AWS API Gateway facade routes one capability at a time, starting with product search, then promotions, while contract tests protect existing mobile clients.
  • New services publish domain events through an outbox to Kafka, but writes to shared tables stop only after each slice owns a schema and reconciliation reaches 99.999%.
  • Running both stacks adds about $45,000 monthly for 12 months, but limits rollback to one capability and avoids freezing 120 developers for a rewrite.

Why interviewers ask this: The interviewer is checking whether strangulation has concrete extraction order, data ownership gates, coexistence cost, and rollback boundaries.

databasedesignlatency

I would deploy two diverse 10 Gbps ExpressRoute circuits and treat hybrid dependencies as temporary migration infrastructure.

  • Circuits terminate in separate peering locations and Azure Virtual WAN hubs, with BGP route summaries preventing 80 workload networks from leaking individual prefixes.
  • Private DNS, Entra ID federation, certificate services, and monitoring are tested before wave one; applications exceeding 15 ms move with their databases in the same wave.
  • The redundant links cost about $32,000 monthly versus $6,000 for VPNs, but VPN throughput and jitter cannot satisfy the database constraint.

Why interviewers ask this: A strong answer derives connectivity from measured bandwidth and latency, includes shared dependencies, and prices redundancy.

deploymentconcurrency

I would choose Apache Kafka because partitioned ordering, replay, and hybrid deployment are primary requirements.

  • Customer ID is the partition key across 384 partitions, sized for under 1,000 events per second each and 30% growth headroom.
  • Three brokers per site with MirrorMaker 2 replicate selected topics, while Schema Registry enforces backward-compatible Avro contracts.
  • Managed Confluent is about $110,000 monthly versus $55,000 self-managed, and I would pay the premium if the team cannot staff three platform engineers for upgrades and recovery.

Why interviewers ask this: The interviewer is evaluating technology selection from ordering and replay constraints plus a realistic operating-cost comparison.

concurrency

I would choose Amazon SQS Standard because the workload needs durable elastic queuing, not broker-side routing or strict ordering.

  • SQS provides visibility timeouts, per-message delay, dead-letter queues, and automatic scaling beyond the 174 jobs per second average and 2,000 per second burst.
  • Workers store job IDs in DynamoDB for idempotency because at-least-once delivery permits duplicates.
  • SQS requests cost roughly $400 monthly, while a highly available RabbitMQ fleet costs about $3,000 plus patching and capacity management.

Why interviewers ask this: A strong answer selects the simpler queue from concrete semantics and quantifies both duplicate handling and price.

dependencies

I would use synchronous calls only for decisions required before checkout and events for work outside that response.

  • Tax uses gRPC with an 80 ms timeout and cached jurisdiction tables because its result changes the payable amount immediately.
  • Fraud uses a parallel gRPC call with a 250 ms budget and a fail-closed rule above $500, while CheckoutCompleted goes to Kafka for analytics.
  • This design spends about $7,000 monthly on Kafka but prevents analytics downtime from consuming the 300 ms checkout budget; making every call synchronous would multiply availability risk.

Why interviewers ask this: The interviewer is checking whether temporal coupling follows business timing rather than one preferred integration style.

distributeddesigntransactions

I would use a Temporal-orchestrated saga with explicit reservation deadlines and compensating actions.

  • The workflow reserves flight and hotel in parallel, authorizes payment last, and persists provider reservation IDs after every successful step.
  • Activity retries use exponential backoff inside the 12-minute window, while cancellation activities are idempotent and continue for 24 hours if a provider is unavailable.
  • Temporal costs about $8,000 monthly at this volume and adds workflow skills, but gives a visible state machine that is safer than hidden Kafka choreography across three external providers.

Why interviewers ask this: A strong answer addresses deadlines, durable state, idempotent compensation, and orchestration cost.

sqlqueriessystem-design

I would use log-based change data capture and publish stable business events through Kafka.

  • Debezium for SQL Server reads the transaction log into raw CDC topics, keeping source overhead below 5% and preserving per-table order.
  • A transformation service converts row changes into CustomerUpdated and InvoicePaid events, removes PII, and uses Schema Registry compatibility checks.
  • The pipeline costs about $14,000 monthly and introduces up to 20 seconds of lag, but removes six polling workloads and prevents source schema from becoming the public contract.

Why interviewers ask this: The interviewer is evaluating low-impact data integration, contract ownership, privacy, lag, and operating cost.

webhooksdesign

I would combine a versioned REST API with durable webhooks and a customer-visible replay endpoint.

  • Apigee validates OAuth 2.0, enforces tenant quotas, and serves OpenAPI v1 with additive evolution and a 12-month retirement window for a future v2.
  • Domain events enter Kafka, webhook workers sign payloads, retry for 24 hours with jitter, and then place failures in a tenant-specific dead-letter store.
  • Apigee and the webhook fleet cost about $48,000 monthly, but managed policy and analytics reduce custom gateway work and support evidence for 3,000 clients.

Why interviewers ask this: A strong answer covers synchronous and asynchronous contracts, replay, tenant isolation, and platform price.

procurementschemadesign

I would land immutable files in Amazon S3 and process them with AWS Glue using vendor-specific contracts and quarantine paths.

  • Vendors upload through S3 Transfer Acceleration or AWS Transfer Family, with manifests, checksums, and object keys partitioned by vendor and business date.
  • Glue jobs validate schemas against the Glue Schema Registry, write Apache Iceberg tables, and send rejected rows to named SQS quarantine queues.
  • About 750 G.2X worker-hours nightly fit near $20,000 monthly; always-on EMR would exceed $35,000 but may win if processing grows beyond the six-hour burst.

Why interviewers ask this: The interviewer is checking batch architecture, schema controls, recoverability, sizing, and technology cost under a deadline.

design

I would use regional S3 data planes with an Iceberg catalog and a separate global metadata control plane.

  • Kinesis Firehose and Kafka Connect write partitioned Parquet, while Lake Formation tags enforce business-unit and data-class access without copying EU payloads outside eu-central-1.
  • Iceberg snapshots support hourly incremental Athena and Spark queries, and compaction targets 512 MB files to avoid millions of small-object scans.
  • The first retained 2 PB costs roughly $47,000 monthly in S3 before compute and grows with retention; replicating it across regions would add at least $40,000 in transfer and violate residency.

Why interviewers ask this: A strong answer combines scale, table format, governance, regional boundaries, and storage economics.

design

I would separate the high-volume telemetry path from the low-latency device command path.

  • AWS IoT Core authenticates devices with individual certificates and routes about 83,000 messages per second into Kinesis using device ID as the partition key.
  • Firehose batches telemetry into S3 and Timestream keeps seven days hot, while MQTT retained topics and device shadows deliver commands within five seconds.
  • IoT messaging and streaming cost about $250,000 monthly before long-term storage and analytics, but operating equivalent global MQTT brokers would require a dedicated team and risk certificate and scaling errors.

Why interviewers ask this: The interviewer is evaluating protocol fit, throughput math, hot-versus-cold storage, command semantics, and managed-service cost.

designapiconcurrency

I would upload directly to object storage and drive transcoding through durable events.

  • The API issues multipart S3 presigned URLs, records upload intent in DynamoDB, and receives an EventBridge Object Created event after completion.
  • EventBridge sends jobs to SQS, AWS MediaConvert creates HLS renditions, and CloudFront serves signed playback URLs; duplicate events are ignored by upload ID.
  • Six PB of new raw video costs about $140,000 for its first month in S3 before renditions and egress, so raw uploads expire after 30 days; self-hosted FFmpeg may cut transcoding cost but adds GPU and codec operations.

Why interviewers ask this: A strong answer keeps large payloads off application servers and prices the managed media trade-off.

Locked questions

  • 21

    Design an e-commerce platform for a launch that rises from 5,000 to 150,000 requests per second in 90 seconds, with checkout inventory accuracy above 99.99%.

    zero-to-onedesign
  • 22

    Design cell-based scaling for 60,000 SaaS tenants and 120,000 requests per second, limiting any single failure to at most 2% of tenants.

    scalingdesigncloud-models
  • 23

    Shard a 30 TB customer database growing by 2 TB monthly when one enterprise tenant generates 18% of all writes and most queries are tenant-scoped.

    databaseshardingqueries
  • 24

    Design caching for a product catalog serving 70,000 reads per second where prices may be stale for 30 seconds but stock availability may be stale for only two seconds.

    designcaching
  • 25

    Design global rate limiting for an API capped at 100,000 requests per second where each customer may briefly exceed its regional share by 5%.

    designcloud-regionsapi
  • 26

    Design disaster recovery for a revenue platform with RTO 15 minutes, RPO one minute, 25 TB of state, and a maximum extra spend of $90,000 monthly.

    designdisaster-recovery
  • 27

    Choose between Kubernetes and serverless for 60 independent APIs, each receiving 5 to 500 requests per second, with a six-person platform team and a $70,000 monthly budget.

    kubernetesapiserverless
  • 28

    Choose a runtime for a licensed risk engine needing 64 GB RAM, local NVMe, startup under 20 seconds, and only 12 production instances.

    compute
  • 29

    Choose a database for 400 million shopping carts, 35,000 writes per second, 30-day TTL, and access only by cart ID.

    database
  • 30

    Choose between Amazon Kinesis and Kafka for 80,000 events per second, 24-hour retention, one AWS region, and three consumer applications.

    kafkacloud-regionsretention
  • 31

    Choose an API gateway for 200 internal services across two clouds, requiring self-hosting, OIDC, mTLS, and less than $250,000 annual license cost.

    gatewayapi-gatewayapi
  • 32

    Decide whether to introduce Istio for 180 Kubernetes services that need mTLS, traffic shifting, and tracing, with a four-person platform team.

    kubernetesmtls
  • 33

    Design a multi-cloud deployment for a regulated service that must survive loss of one provider, processes 2,000 transactions per second, and has a $300,000 monthly ceiling.

    designmulti-cloudtransactions
  • 34

    Choose build or buy for customer identity supporting 8 million users, social login, MFA, EU residency, and a launch in six months.

    decision-making
  • 35

    Design observability for 900 services producing 25 TB of logs daily, with 30-day searchable retention and a $220,000 monthly cap.

    retentiondesignobservability
  • 36

    Design Zero Trust access for 14,000 employees, 2,000 contractors, and 600 workloads across AWS, Azure, and on-premises systems.

    system-designdesignzero-trust
  • 37

    Design PCI DSS scope reduction for a marketplace processing 20 million card transactions monthly without storing primary account numbers in product services.

    transactionsdesignconcurrency
  • 38

    Design a SOC 2-ready AWS landing zone for 120 accounts that must onboard a team in under one hour and collect evidence continuously.

    designlanding-zoneonboarding
  • 39

    Design an AWS account structure for 30 product teams, four environments, and six regulated workloads without creating hundreds of unnecessary network connections.

    design
  • 40

    Design network connectivity for 300 VPCs across three regions, 40 on-premises sites, overlapping address ranges, and centralized egress inspection.

    designcloud-regions
  • 41

    Reduce a $180,000 monthly cloud bill where $65,000 is inter-region and internet egress, while keeping p95 latency below 100 ms for global users.

    cloud-regionslatency
  • 42

    Plan compute purchasing for a platform with a steady 12,000 vCPU baseline, nightly growth to 20,000, and unpredictable monthly peaks of 35,000.

  • 43

    Redesign an image API serving 4 billion requests monthly whose serverless bill rose to $260,000, with p99 required below 180 ms.

    serverlessapi
  • 44

    Decide whether to use Event Sourcing for an insurance claims system processing 80,000 claims daily with seven-year audit and point-in-time reconstruction requirements.

    event-sourcingsystem-designconcurrency
  • 45

    Decide whether to apply CQRS to a catalog with 5,000 writes per second, 120,000 reads per second, and six distinct search views refreshed within 10 seconds.

    distinctcqrs
  • 46

    Choose a search architecture for 600 million product documents, 20,000 updates per second, multilingual queries, and p95 below 120 ms.

    queriesarchitecture
  • 47

    Integrate Salesforce with an order platform at 10,000 updates per minute, with two-minute freshness and no direct database access on either side.

    database
  • 48

    Design B2B file exchange for 900 partners sending 30,000 files daily, from 10 MB to 20 GB, with four-hour processing and seven-year audit retention.

    retentiondesignconcurrency
  • 49

    Expose a legacy SOAP billing system to 40 new REST consumers within five months without letting its 60 requests per second limit cascade upstream.

    system-designrest
  • 50

    Plan integration of an acquired company with 70 applications, a different identity provider, and 6 TB of customer data within 12 months while both businesses keep operating.

  • 51

    Four hours after moving a 6 TB billing database to Aurora PostgreSQL, you find 18,000 writes split between Oracle and Aurora. What do you do before reopening traffic?

    databasepostgres
  • 52

    A planned 20-minute CRM cutover has reached 55 minutes, CDC lag is still 14 minutes, and the business allows at most 60 minutes of downtime. What do you do?

  • 53

    After an SAP integration release, 7,400 duplicate InvoiceCreated events produce $2.1 million in duplicate receivables. How do you contain and correct it?

  • 54

    During a 12 TB SQL Server migration, Debezium falls six hours behind and 3% of customer balances differ in the target. How do you decide whether to roll back?

    sqlmigrationsrollback
  • 55

    Your CIAM vendor raises its price from $70,000 to $119,000 per month with 90 days' notice, and eight million users depend on it. What is your exit response?

    procurement
  • 56

    A document vendor cuts your API quota from 10,000 to 2,000 requests per minute two weeks before tax season, while peak demand is 7,500. What do you change?

    procurementapi
  • 57

    A payroll SaaS scheduled for termination allows exports at only 50 GB per day, but you must move 9 TB in 120 days and retain seven years of audit history. What do you do?

    cloud-modelscloud
  • 58

    A proprietary Kubernetes control plane is discontinued with six months' support remaining, and 140 services use its custom deployment API. How do you exit without freezing releases?

    kubernetesdeploymentapi
  • 59

    Three teams reject your Kafka proposal after a pilot adds 11 minutes of processing lag against a two-minute requirement. Do you defend or change the architecture?

    kafkaconcurrency
  • 60

    Security rejects your public API architecture because a threat test found 14 cross-tenant data leaks in 100,000 requests, delaying a launch by three weeks. What do you do?

    api
  • 61

    Finance opposes your $160,000 monthly active-active design after a chaos test shows warm standby recovers in 18 minutes against a 30-minute RTO. What do you recommend now?

    design
  • 62

    Sales and support reject your proposed Customer domain split because it adds three APIs and threatens a launch in 10 weeks. How do you defend or revise it?

    api
  • 63

    A $1.2 million modernization budget is cut to $700,000 after contracts are signed, with the first customer release due in five months. What do you cut?

  • 64

    Your approved cloud run budget falls from $180,000 to $110,000 per month, but the platform still owes 99.95% availability to 20,000 tenants. How do you replan?

  • 65

    A migration program planned at $3 million is reforecast at $4.2 million after only 30 of 90 applications move. How do you replan the remaining nine months?

    migrationscloud-migration
  • 66

    An observability design budgeted at $220,000 per month is capped at $120,000 after 900 services already emit 25 TB of logs daily. What do you change?

    designobservability
  • 67

    Ten weeks before launch, a PCI review finds card numbers in Kafka logs across 14 services, and the deadline cannot move. What do you do?

    estimationkafka
  • 68

    A regulator gives you 16 weeks to keep EU records in-region, but a 20-year-old mainframe sends full customer files to a US batch processor nightly. What do you change?

    batchcloud-regionsconcurrency
  • 69

    A SOC 2 auditor gives you 30 days to eliminate 46 unmanaged production changes across 120 AWS accounts. What do you do without freezing all delivery?

  • 70

    A partner mandates TLS 1.3 in six weeks, but 280 legacy clients only support TLS 1.2 and generate 35% of revenue. How do you meet the constraint?

    tls
  • 71

    After a release, checkout p99 rises from 220 ms to 1.4 seconds at 8,000 requests per second. What do you do in the first hour and after stabilization?

  • 72

    A single availability-zone failure takes down a service promised at 99.99% for 47 minutes because all pods used one subnet. What concrete fix do you make?

    networkingpromises
  • 73

    Kafka consumer lag jumps from 20 seconds to 38 minutes during a 120,000-event-per-second campaign, missing a five-minute order-completion SLO. What do you change?

    kafkaslo
  • 74

    A regional database failover takes 42 minutes against a 15-minute RTO because Terraform capacity quotas are missing. What do you change before the next drill?

    capacitycloud-regionsfailover
  • 75

    A Redis invalidation failure shows out-of-stock inventory for 11 minutes and causes 3,200 oversold orders. What do you do?

    redis
  • 76

    A nightly 9 TB vendor load finishes in 9.5 hours instead of the six-hour contractual window after file volume grows 40%. What do you change?

    procurement
  • 77

    A junior architect proposes Istio, Kafka, and CQRS for a service handling 40 requests per second, adding six weeks to a 10-week delivery. How do you mentor them?

    cqrskafkamentoring
  • 78

    An architect you mentor schedules a 4 TB database cutover with no rollback rehearsal and only a 30-minute outage window. What do you do?

    mentoringdatabaserollback
  • 79

    Two architects on your team produce opposite API designs and have blocked 18 engineers for nine days. How do you mentor them to close the decision?

    mentoringapidesign
  • 80

    A senior architect has underestimated cloud cost by more than 30% on three consecutive proposals. How do you mentor them before the next $2 million bid?

    mentoring
  • 81

    Two days after federating an acquired company's identity provider, 1,800 users are locked out and 63 former employees regain access. What do you do?

  • 82

    An API release breaks 26 of 180 partners and causes a 70-minute order outage despite a two-year compatibility promise. What do you do?

    promises
  • 83

    An ERP vendor unexpectedly limits writes to 60 requests per second, while your new integration sends 240 and is dropping 8% of orders. What do you change?

    procurement
  • 84

    After moving fulfillment to events, 4,600 orders execute out of sequence and 900 ship before payment capture. How do you repair the architecture?

    architecture
  • 85

    Operations opposes your active-active design after a game day shows 37 minutes of split-brain writes and 12,000 conflicts. Do you defend or change it?

    design
  • 86

    The board insists on two-cloud active-active in six months, but a four-week prototype reaches only 97.8% consistency and costs $410,000 per month against a $300,000 cap. What do you recommend?

    consistencyprototypes
  • 87

    Product teams reject a mandatory API gateway because the pilot adds 28 ms p99 against a 15 ms budget across 200 services. What do you do?

    gatewayapi-gatewayapi
  • 88

    A serverless platform you recommended suffers 9% cold-start timeouts during a 30-times burst, and the owning team wants Kubernetes instead. How do you respond?

    resilienceserverlesskubernetes
  • 89

    A data-locality mistake drives inter-region egress from $40,000 to $105,000 per month, but Finance demands a fix within 30 days without raising p95 above 100 ms. What do you do?

    ownershipcloud-regions
  • 90

    A banking app must launch in eight weeks under a residency rule, but legal approval forbids cross-region backups and the contract promises 99.99% availability. What do you do?

    backupscloud-regionspromises
  • 91

    An expired certificate causes a 96-minute outage across 300 legacy applications, but 70 cannot use your automated service identity system. What do you fix first?

    system-design
  • 92

    A disaster recovery drill restores only 21 TB of a 25 TB platform and misses the 15-minute RTO by 34 minutes. What do you tell executives and fix?

    disaster-recovery
  • 93

    AWS Config finds 380 Terraform-managed resources drifted across 45 accounts one week before a compliance audit. What do you do?

    terraformconfigiac
  • 94

    Ransomware encrypts 600 virtual machines, and the only clean legacy backup is 18 hours old against a four-hour RPO. What do you do under pressure?

    backupscomputeencryption
  • 95

    Sales promises a 99.99% SLA and 30-minute regional recovery for a launch in six weeks, but the latest drill takes 74 minutes. What do you do before signature?

    cloud-regionspromises
  • 96

    After a failed migration, an architect you mentor blames Operations, although logs show their runbook omitted a 12-minute replication check. How do you respond?

    migrationsreplicationmentoring
  • 97

    Runtime telemetry shows 17 teams bypassed the approved API layer and created 43 unrecorded data flows before a release in five days. What do you do?

    api
  • 98

    A proprietary database vendor raises licenses by 80%, taking annual cost from $1.5 million to $2.7 million, and your exit test finds 140 stored procedures with no PostgreSQL equivalent. What do you do?

    databasepostgresstored-procedures
  • 99

    Users in APAC see 620 ms p99 against a 200 ms target after a data-residency change forces writes to Frankfurt. What concrete fix do you make?

  • 100

    A $30 million program is blocked for 21 days because four architects keep reopening the integration decision, and the first release is due in eight weeks. What do you do?