Skip to content

Azure Engineer interview questions

100 real questions with model answers and explanations for Senior Azure Engineer candidates.

See a Azure Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

I would organize management groups around stable governance boundaries, not the company reporting chart.

  • Place Platform, Landing Zones, Sandbox, and Decommissioned under the tenant root so lifecycle controls have clear inheritance.
  • Split Landing Zones into Corp, Online, and Regulated branches, then separate production from nonproduction where policy must differ.
  • Assign subscriptions to one branch through automation and keep workloads out of the tenant root and Platform hierarchy.
  • Pilot every hierarchy or inherited-policy change on 2 canary subscriptions before rolling it across all 80.

Why interviewers ask this: The interviewer is evaluating management-group inheritance, subscription boundaries, and controlled change at enterprise scale.

slo

I would isolate connectivity, management, identity, and security operations in dedicated platform subscriptions.

  • Connectivity owns regional hubs, ExpressRoute, Private DNS Resolver, Azure Firewall, and DDoS plans without granting workload access.
  • Management owns Azure Monitor workspaces, Automation, backup governance, and shared operational tooling for the 99.95% SLO.
  • Identity contains only identity-dependent services such as domain controllers when they are required, while Entra remains tenant-scoped.
  • Security owns Sentinel and Defender administration, with PIM roles distinct from the network team and workload owners.

Why interviewers ask this: The interviewer is checking whether shared services and privileged duties are separated without inventing unnecessary subscriptions.

designerror-handling

I would publish versioned initiatives at management-group scope and make every exception explicit and expiring.

  • Separate security, diagnostics, region, tagging, and network baselines so each initiative has one owner and release cadence.
  • Use Deny only for proven invariants, Audit first for new controls, and DeployIfNotExists with managed remediation identities for missing settings.
  • Store exemptions with owner, justification, affected resource scope, and a maximum 30-day expiry instead of cloning assignments.
  • Report compliance through Policy Insights and create remediation tasks until each branch reaches the 95% target within 24 hours.

Why interviewers ask this: The interviewer is testing policy composition, remediation, exception governance, and measurable compliance.

activation

I would grant Azure access through Entra groups and make privileged roles eligible through PIM rather than permanently active.

  • Map job functions to groups for Reader, Contributor, Network Operator, and narrowly scoped custom roles, avoiding direct user assignments.
  • Require phishing-resistant MFA, approval, ticket reference, and a 1-hour activation for Owner, User Access Administrator, and platform roles.
  • Keep 2 cloud-only break-glass accounts outside Conditional Access, protect them with hardware credentials, and test them quarterly.
  • Run quarterly access reviews and alert on permanent privileged assignments, role changes, and activations outside approved hours.

Why interviewers ask this: The interviewer is evaluating scalable authorization, just-in-time privilege, and a controlled emergency path.

I would expose a validated request contract backed by an idempotent subscription-vending workflow.

  • Capture owner group, billing scope, management group, environment, data class, region, address space, and monthly budget before creation.
  • Create the subscription through a subscription alias, move it to the target management group, and wait for inherited Policy assignments.
  • Deploy the baseline for RBAC, budgets, diagnostics, Defender, networking, and resource-provider registration from one versioned release.
  • Track request-to-ready p95 against 30 minutes and reconcile by request ID so retries never create duplicate subscriptions.

Why interviewers ask this: The interviewer wants a concrete self-service factory that handles billing, governance, networking, and retry safety.

capacity

I would manage provider registration, quotas, commercial reservations, and physical capacity as independent controls.

  • Grant the vending identity RBAC permission for providers/register/action only where needed; Azure Policy cannot block namespace registration, but it can deny subsequent resource types or configurations.
  • Inventory regional vCPU, public IP, gateway, AKS, and service-specific quotas in all 15 subscriptions, then submit increases at least 4 weeks early.
  • Create On-demand Capacity Reservations for the required VM families and zones because they cover capacity, while Azure Reservations discount eligible usage but do not reserve capacity.
  • Rehearse deployment at the 12,000-vCPU scale and validate provider state, quota, and deployability separately in both the primary and fallback regions.

Why interviewers ask this: The interviewer is checking RBAC-based provider registration, Policy's actual enforcement boundary, quota, cost discounts, and capacity assurance.

I would choose secured Virtual WAN because the global branch and VNet scale makes managed routing more valuable than custom hub control.

  • Build one secured virtual hub per region and segment production, nonproduction, shared services, and regulated traffic through route tables.
  • Connect the 25 branches with VPN or ExpressRoute and use hub routing intent to steer private and internet traffic through Azure Firewall.
  • Manage connectivity through Virtual WAN and Azure Virtual Network Manager rather than maintaining hundreds of bilateral peerings.
  • I would retain hub-spoke for 2 regions and fewer than 40 VNets when appliance choice and per-route control outweigh operating simplicity.

Why interviewers ask this: The interviewer is evaluating whether topology follows network scale, branch connectivity, and the operating model.

hybrid-networking

I would remove shared physical and routing dependencies across both ExpressRoute paths and keep VPN as a lower-preference route.

  • Terminate circuits in different peering locations and on separate customer routers, then use zone-redundant ExpressRoute gateways.
  • Advertise summarized prefixes with BGP and set path preference deliberately so stateful inspection does not see asymmetric flows.
  • Build active-active VPN gateways over 2 internet providers and advertise the same prefixes with less-preferred BGP attributes.
  • Withdraw each circuit, peering site, router, and gateway quarterly while measuring convergence against the 90-second target.

Why interviewers ask this: The interviewer is testing physical diversity, BGP behavior, gateway resilience, and measurable fallback.

designdnsprivate-connectivity

I would centralize hybrid resolution with Azure DNS Private Resolver while keeping authoritative private zones centrally governed.

  • Deploy inbound and outbound endpoints in dedicated subnets in each required region and link forwarding rulesets to authorized VNets; Private Resolver is built-in zone-redundant, so endpoints are not manually assigned to 2 availability zones.
  • Forward only the 2 on-premises suffixes outbound and configure on-premises conditional forwarders for Azure private namespaces.
  • Link each of the 40 Private DNS zones through automation, avoiding duplicate same-name zones that create inconsistent answers.
  • Set record TTLs and deployment checks to meet 5 minutes, then monitor query failures, NXDOMAIN rates, and forwarding loops.

Why interviewers ask this: The interviewer is checking hybrid DNS direction, built-in resolver resilience, zone ownership, VNet links, and propagation constraints.

dnsprivate-connectivitycloud

I would make private endpoints and their DNS records part of the service deployment contract, then deny public access.

  • Create endpoints in dedicated spoke subnets sized for 25 services and apply least-privilege NSG rules where private-endpoint network policies are enabled.
  • Use private DNS zone groups so endpoint creation registers records in centrally owned zones without manual A records.
  • Resolve zones through Private DNS Resolver for both Azure and on-premises clients, testing each service's required subresources.
  • Apply Policy to deny public network access and audit orphaned endpoints, stale DNS records, and unauthorized cross-subscription approvals.

Why interviewers ask this: The interviewer is evaluating Private Link lifecycle, DNS integration, and enforcement of a no-public-endpoint constraint.

I would deploy a regional secured hub with Azure Firewall Premium so egress remains regional and inspectable.

  • Route spoke defaults through the local firewall by routing intent or UDRs, and use Azure Firewall Manager for hierarchical policy.
  • Enable DNS proxy, TLS inspection only for approved categories, and threat-intelligence blocking with documented bypasses for pinned traffic.
  • Use multiple public IPs or a NAT Gateway association where supported to provide enough SNAT ports beyond the 20 Gbps load profile.
  • Deny workload public IPs and unapproved route tables with Policy, then alert on firewall throughput, SNAT use, and bypass routes.

Why interviewers ask this: The interviewer is checking regional autonomy, centralized policy, SNAT capacity, and prevention of alternate egress.

designedge-routingapi

I would use Front Door Premium with active-active regional origins reached through Private Link.

  • Put each regional API behind an internal origin, enable Private Link approval, and reject traffic that bypasses Front Door.
  • Configure latency routing, health probes on dependency-aware endpoints, and enough regional capacity to absorb one region's 30,000 requests per second.
  • Apply WAF managed rules, rate limits, bot controls, and certificate rotation at the edge without caching sensitive API responses.
  • Keep sessions and data external to compute, then test regional origin removal quarterly against the 5-minute RTO and 99.99% SLA.

Why interviewers ask this: The interviewer is evaluating edge routing, private origins, regional capacity, security, and tested recovery.

containerskuberneteszero-to-one

I would place each workload by execution constraint and make the simplest managed platform the default.

  • Use Functions for short event handlers and Container Apps for bursty containers that need KEDA scaling but not Kubernetes APIs.
  • Use App Service for conventional web APIs needing slots and managed runtime operations, and VMs for kernel, appliance, or licensing constraints.
  • Choose AKS only when several teams require Kubernetes scheduling, controllers, network policy, or a shared container platform.
  • Benchmark p99 latency, cold starts, scale limits, staffing, and monthly cost at 10,000 requests per second before standardizing.

Why interviewers ask this: The interviewer is checking whether compute choice follows workload constraints rather than a single preferred platform.

procurementdesign

I would use a zone-spanning Virtual Machine Scale Set built from an immutable, tested image.

  • Produce the Windows image with Azure VM Image Builder, pin the vendor version, and scan it before Shared Image Gallery replication.
  • Spread at least 2 instances across zones behind Standard Load Balancer or Application Gateway and reserve the required 64-vCPU SKU capacity.
  • Apply Azure Update Manager in staged rings and keep application data on a managed service or replicated disks supported by the vendor.
  • Autoscale only within licensing limits and prove zone loss plus image rollback against the 99.95% requirement.

Why interviewers ask this: The interviewer is evaluating when VMs are justified and how images, zones, licensing, and patching are controlled.

deploymentrollbackkubernetes

I would use zone-redundant Premium plans with private endpoints, VNet integration, and deployment slots.

  • Group APIs only when their scaling and blast-radius profiles match, keeping critical services on separate plans across 3 instances.
  • Route inbound traffic through Application Gateway or Front Door to private endpoints and send outbound calls through VNet integration.
  • Warm a staging slot, run dependency checks, then swap with production while retaining the previous slot for a sub-10-minute rollback.
  • Use managed identities, Key Vault references, autoscale rules, and per-app telemetry instead of shared credentials or manual configuration.

Why interviewers ask this: The interviewer is checking App Service isolation, private networking, slot-based delivery, and rollback design.

eventscontainers

I would use Functions for short event handlers and Container Apps jobs or services for long-running containerized work.

  • Put the 200-millisecond handlers on Functions with event triggers, idempotency, and a hosting plan chosen from measured cold-start limits.
  • Run 45-minute work as Container Apps jobs with explicit CPU, memory, retry, timeout, and parallelism settings.
  • Use KEDA-backed scaling from Service Bus queue depth while capping replicas so a 30-fold burst cannot exhaust databases.
  • Separate poison messages, emit correlation IDs, and compare idle plus burst cost before selecting Consumption or dedicated capacity.

Why interviewers ask this: The interviewer is evaluating duration limits, scaling behavior, downstream protection, and cost across serverless options.

designkubernetes

I would separate security isolation from availability: shared clusters can host soft tenants, while regulated trust boundaries get dedicated clusters or subscriptions.

  • Give each team namespaces with ResourceQuota, LimitRange, network policy, Pod Security admission, and Entra-backed RBAC.
  • Use Azure Workload Identity per service account so the 80 services never share cluster or cloud credentials.
  • Run production clusters on the supported Standard or Premium tier and spread system and user node pools across availability zones with capacity for a zone loss.
  • Deploy each service with multiple replicas, topology spread, PodDisruptionBudgets, and tested dependency failover; 2 clusters reduce blast radius but do not by themselves prove a 99.95% workload SLA.

Why interviewers ask this: The interviewer is checking security isolation, supported AKS tier selection, zonal node design, and workload-level availability controls.

designresilienceservice-mesh

I would separate stable, burst, and Spot capacity into node pools and validate the mesh with a representative load test.

  • Use cluster autoscaler or node auto-provisioning with pod requests, topology spread, and disruption budgets that preserve critical replicas.
  • Taint Spot pools and schedule only checkpointed, retryable jobs there, with regular nodes available before Azure eviction completes.
  • Adopt the AKS Istio add-on only for required mTLS and traffic policy, measuring sidecar CPU, memory, and p99 against 10 milliseconds.
  • Set scaling ceilings from subnet IPs, regional vCPU quota, database capacity, and the 70% demand swing rather than CPU alone.

Why interviewers ask this: The interviewer is evaluating autoscaling boundaries, Spot interruption safety, and evidence-based service-mesh adoption.

deploymentrollbackkubernetes

I would publish a versioned service template and make progressive delivery the default path, not a bespoke team choice.

  • The template creates workload identity, probes, resource requests, telemetry, policy checks, dashboards, and an Azure DevOps or GitHub pipeline.
  • Build once, sign the image in ACR, promote the digest through environments, and reconcile deployments with Flux.
  • Release through 5%, 25%, and 50% canary steps or blue-green for incompatible changes, gating on error rate and p99 latency.
  • Keep the previous manifest and image digest deployable so an automated abort restores service within 10 minutes.

Why interviewers ask this: The interviewer is checking whether platform standards improve delivery while preserving measurable rollback safety.

sqldesign

I would challenge the RPO because an Azure SQL failover group alone cannot guarantee less than 30 seconds of data loss.

  • Enable zone redundancy on the primary and secondary databases or elastic pools, not on their logical servers, and size both sides for 5,000 TPS.
  • Treat geo-replication as asynchronous: renegotiate the hard RPO or add and test a business-layer synchronous design with idempotency, ordering, and write fencing.
  • Use the read-write listener for writes, restrict the read-only listener to staleness-tolerant reads, and monitor replication lag plus dependency health.
  • Use monitored manual failover because the automatic grace-period minimum is 1 hour, and rehearse promotion quarterly against the 5-minute RTO and agreed RPO.

Why interviewers ask this: The interviewer is evaluating asynchronous RPO limits, database-level zone redundancy, listener behavior, and a realistic manual failover contract.

Locked questions

  • 21

    A SaaS platform has 40 Azure SQL tenant databases, requires coordinated regional recovery within 20 minutes, and permits 2 minutes of data loss; how would you organize failover?

    sqldatabasecloud
  • 22

    Which Azure Cosmos DB consistency model would you choose for 20,000 requests per second across 3 regions when profile reads may be 5 seconds stale but balance reads may not be stale?

    cosmos-dbconsistency
  • 23

    A 3-region Cosmos DB account accepts multi-region writes for 10,000 orders per second, and conflicts must be resolved within 60 seconds without losing updates; what conflict strategy would you use?

    cosmos-db
  • 24

    How would you process a Cosmos DB change feed carrying 50 million changes per day when duplicates are allowed but no business event may be lost for more than 15 minutes?

    cosmos-dbconcurrency
  • 25

    A regulated archive stores 100 TB of blobs for 7 years, requires WORM retention, and needs a second Azure region with RPO under 15 minutes; how would you design it?

    designregionsretention
  • 26

    How would you design Redis caching for 100,000 operations per second with p99 under 5 milliseconds and no more than 60 seconds of stale catalog data?

    redisdesigncaching
  • 27

    An order workflow emits 20 million events per day and must converge across 6 services within 2 minutes without distributed transactions; how would you design consistency?

    distributeddesignconsistency
  • 28

    How would you design Conditional Access and privileged access workstations for 2,000 users and administrators across 80 Azure subscriptions, including risk-based authentication and 2 emergency accounts?

    authdesign
  • 29

    Three hundred Azure workloads must access storage, SQL, and Key Vault with zero client secrets and credential rotation under 24 hours; how would you use managed identities?

    secretssql
  • 30

    A payment platform manages 500 keys, requires FIPS-protected HSM keys, 90-day rotation, and recovery from deletion for 90 days; would you choose Key Vault Premium or Managed HSM?

    secrets
  • 31

    How would you combine Defender for Cloud and Microsoft Sentinel across 80 subscriptions producing 2 TB of security data per day with a monthly ingestion cap of $150,000?

  • 32

    An audit requires 98% Azure Policy compliance across 80 subscriptions within 30 days, but 12 legacy applications cannot meet 3 controls; how would you close the gap?

  • 33

    How would you retain immutable Azure Activity and diagnostic logs for 7 years across 80 subscriptions while keeping only 90 days searchable in Log Analytics?

    immutability
  • 34

    A software platform ships 200 container images to 4 countries, requires signed artifacts, and forbids customer data or encryption keys from leaving each country; how would you secure the supply chain and residency boundary?

    encryptionsupply-chaincontainers
  • 35

    Would you standardize on Bicep, Terraform, or authored ARM templates for 80 Azure subscriptions and 12 teams when 90% of resources are Azure-native?

    terraformiac
  • 36

    How would you partition workload Terraform state for 80 subscriptions and 30 concurrent pipelines so corruption affects at most 1 subscription, while shared management-group and Policy foundation state remains a separate higher-blast-radius exception?

    terraformci-cdconcurrency
  • 37

    How would you run a Bicep module registry containing 40 modules for 12 teams when critical fixes must reach production within 7 days without breaking existing deployments?

    deploymentregistriesiac
  • 38

    How would you implement policy as code for 60 Azure Policy definitions across 80 subscriptions when every change needs 2 approvals and a 14-day canary period?

    deployment-strategies
  • 39

    Eighty subscriptions must report infrastructure drift within 24 hours, but automatic correction may not restart production resources; how would you detect and handle drift?

    iac
  • 40

    How would you roll out a landing-zone change across 80 subscriptions in 10 waves when rollback must complete within 30 minutes and no wave may exceed 10 subscriptions?

    rollback
  • 41

    How would you define and monitor a 99.95% monthly availability SLO for 25 Azure APIs when the error budget is about 21.6 minutes?

    reliabilityslomonitoring
  • 42

    How would you instrument 150 services producing 100,000 spans per second with OpenTelemetry and Azure Monitor while keeping telemetry spend below $80,000 per month?

    monitoringobservability
  • 43

    Forty applications need 3 disaster-recovery tiers with RTO/RPO pairs of 15 minutes/1 minute, 4 hours/1 hour, and 24 hours/24 hours; how would you map Azure services and cost?

  • 44

    How would you protect backups for 300 VMs and 80 TB of data from ransomware when retention is 1 year and restore evidence is required every 90 days?

    retentionbackups
  • 45

    Two hundred on-premises VMs need Azure Site Recovery with RTO under 30 minutes and RPO under 15 minutes; how would you design and test the recovery plan?

    recoverydesign
  • 46

    How would you run an Azure Well-Architected review for a platform spending $2 million per year with a 99.99% SLA and 6 weeks to fund the top 10 improvements?

  • 47

    An Azure estate spends $4 million per year and 65% of compute usage is stable; how would you choose reservations versus Azure Savings Plan without overcommitting for 3 years?

  • 48

    How would you allocate at least 95% of Azure cost across 80 subscriptions and 15 business units when shared network and security services represent 12% of monthly spend?

  • 49

    You must cut a $500,000 monthly Azure bill by 25% within 6 months without reducing a 99.95% SLA; what concrete FinOps plan would you execute?

    finops
  • 50

    Twelve teams disagree on an Azure platform standard that affects 80 services, and 8 engineers must deliver an RFC within 4 weeks under a $200,000 migration cap; how would you lead the decision?

    conflictdecision-makingmigrations
  • 51

    Azure Front Door sends 65% of 24,000 requests per second to West Europe; 4 external probes show 31% checkout failures for 8 minutes after a WAF policy change, while origin health remains green. How do you contain the incident, and what evidence gates traffic recovery?

    incidentsedge-routing
  • 52

    An Azure Virtual WAN route-table association leaked 10.42.0.0/16 from development into 180 production VNets; Network Watcher flow logs show 16,000 denied flows and 7 payment services have been unreachable for 11 minutes. What do you contain first, and what proves the routes are safe?

  • 53

    A 10 Gbps ExpressRoute circuit failed 14 minutes ago and traffic moved to two 2 Gbps site-to-site VPN tunnels, but packet captures show return traffic still preferring ExpressRoute and 29% of API calls time out at 3.6 Gbps. What do you change, and what recovery gate prevents asymmetry?

    hybrid-networkingapi
  • 54

    Azure NAT Gateway serves 52,000 concurrent outbound connections; SNAT connection failures reach 8,600 per minute and partner TLS errors hit 26%, while flow logs identify one crawler opening 45 connections per second per replica. How do you contain and gate recovery?

    gatewaynetworkingreplication
  • 55

    Application Gateway v2 returns 502 for 18% of 42,000 requests per second after a backend release; access logs show ERRORINFO_UPSTREAM_CONNECTION_RESET and UnhealthyHostCount rose from 2 to 94, while backend CPU is 61%. Do you scale, roll back, or change the gateway, and what gates recovery?

    gatewayload-balancingrollback
  • 56

    After an Azure Private DNS Resolver ruleset update, 68 VNets cannot resolve corp.internal for 12 minutes; DNS query logs show 84% SERVFAIL, both outbound endpoints are provisioned, and on-premises DNS answers direct queries in 14 ms. What do you roll back, and what proves recovery?

    queriesdnsendpoints
  • 57

    A storage account Private Endpoint was replaced 20 minutes ago; 140 consumers still resolve the old private IP, 38 new connections are pending approval, and application logs show 33% TCP timeouts. What do you restore, and how do you gate the Private Link migration?

    resiliencenetworkingprivate-connectivity
  • 58

    A Front Door origin change sent 88 TB per day of EU-regulated traffic through East US for 29 minutes; flow logs show 34 Gbps crossing the residency boundary. How do you contain the breach, preserve production evidence, and gate compliant recovery?

    edge-routing
  • 59

    A VM Scale Set rolling upgrade put a bad image on 170 of 240 instances just as 65% of Spot capacity was evicted; boot diagnostics show cloud-init exit 1, healthy instances fell to 52, and API errors reached 24% in 7 minutes. How do you recover and gate replacement?

    capacityautoscalingapi
  • 60

    An App Service deployment-slot swap completed for 60 instances, but checkout errors rose from 0.3% to 13% while health checks stayed green; Application Insights shows the production slot using a staging payment endpoint for 41% of calls. What do you swap back, and what gates another attempt?

    deploymenthealth-checksendpoints
  • 61

    A new Azure Container Apps revision receives 50% of 36,000 requests per second; 5xx reaches 16%, replicas are ready, and Log Analytics ties failures to revision 2026-07-16-2. What do you contain, and what recovery gate was missing?

    replicationcontainers
  • 62

    An AKS cluster with 780 nodes shows Kubernetes API p99 of 13 seconds, 34 nodes NotReady, and 21% failed requests 6 minutes after an Azure CNI upgrade; Azure Service Health reports no control-plane incident. Do you replace nodes, roll back CNI, or fail over, and what proves recovery?

    kubernetesrollbackapi
  • 63

    The AKS cluster autoscaler removed 120 nodes in 9 minutes after a node-pool limit change; 2,100 pods are Pending, payment p99 is 4.1 seconds, and autoscaler logs cite unsatisfied topology constraints. What do you stop, and how do you gate scheduling recovery?

    jobsscalingkubernetes
  • 64

    After an Istio sidecar update on 82 AKS services, p99 rose from 190 ms to 2.1 seconds and retries amplified traffic from 28,000 to 75,000 requests per second; traces show 690 ms in Envoy while application CPU is 44%. Do you tune or roll back, and what gates recovery?

    rollbackkubernetes
  • 65

    A retry bug drives Azure Functions from 7,000 to 49,000 executions per second; host concurrency reaches 92,000, 12 unrelated apps are throttled, and Service Bus oldest-message age reaches 38 minutes. What concurrency decision do you make, and how do you drain safely?

    resilienceserverlessmessaging
  • 66

    Microsoft Defender for Cloud flags a critical RCE in an ACR base image used by 1,600 running containers across 24 services; no exploitation is confirmed, but the patched digest fails 8% of canary health checks. Do you stop production, accept risk, or roll back, and what gates replacement?

    containershealth-checksrollback
  • 67

    Azure SQL Database failed over in 42 seconds, but 1,100 application instances reconnected at once; sessions reached 14,700 of 15,000, CPU is 95%, and checkout errors remain 18% after 6 minutes. How do you stop the storm and gate writer recovery?

    sqldatabasesessions
  • 68

    Azure SQL Managed Instance has 110 GB free on 11 TB, storage falls 19 GB per hour, replica lag is 24 minutes, and Query Store ties growth to an index operation started 4 hours ago. Do you add storage, cancel, or fail over, and what gates recovery?

    sqlindexesqueries
  • 69

    An Azure SQL readable secondary serving 64% of catalog reads is 43 minutes behind while the primary handles 20,000 writes per second; stale prices cause 7% cart mismatches and Data IO reaches 98%. Do you promote, stop reads, or rebuild, and what gates return?

    sql
  • 70

    A Cosmos DB container receives 180,000 requests per second, but one tenant owns 46% of traffic; normalized RU consumption for its partition hits 100%, 429 responses reach 22%, and other tenants time out. How do you contain the hot key and gate recovery?

    normalizationpartitioningcosmos-db
  • 71

    A Cosmos DB multi-write account accepted updates in 2 regions during a 9-minute partition; conflict feed contains 74,000 customer-profile conflicts and 6% of loyalty balances disagree. What do you fence, and what evidence gates multi-region writes?

    partitioningcosmos-dbconflict
  • 72

    A lifecycle rule deleted 8.4 million source blobs in 17 minutes; object replication had completed for 6.3 million, while 2.1 million show failed or missing replication status and soft delete is enabled for 14 days. How do you contain and recover without losing the remaining copy?

    replicationsoft-delete
  • 73

    Azure Managed Redis failed over in 76 seconds; 900 application instances then missed cache simultaneously, database CPU reached 97%, and checkout p99 hit 5.2 seconds. How do you contain the stampede and gate cache recovery?

    databaserediscaching
  • 74

    Service Bus backlog reaches 11 million messages and Event Hubs consumer lag reaches 38 minutes after a downstream release; dead-letter rate is 9%, SQL CPU is 88%, and duplicate invoices appear. What do you pause, and how do you gate backlog recovery?

    messagingstreamingsql
  • 75

    Microsoft Entra sign-in logs show a service principal authenticating from 3 countries and creating 9 credentials in 14 minutes; it made 4,200 Key Vault reads and Defender alerts indicate token replay. How do you contain the compromise and gate identity recovery?

    tokenssecretsalerting
  • 76

    PIM automation removed eligible Owner role schedules from 46 subscriptions; 18 responders cannot activate emergency roles during a production outage, and PIM audit history plus Activity Logs show its managed identity deleted 46 schedules 9 minutes ago. How do you restore access without creating a permanent backdoor?

    identity
  • 77

    A Key Vault key was disabled during rotation; 23 applications now fail 61% of decrypt operations, Key Vault audit logs show key version 7 disabled 6 minutes ago, and 18 million records still reference it. What do you restore, and what gates rekeying?

    secrets
  • 78

    A database secret rotation changed Key Vault 12 minutes before all 320 application replicas reloaded it; authentication failures reached 37%, while 140 replicas still use the old password. How do you contain and gate a safe rotation?

    authsecretspasswords
  • 79

    Defender for Cloud detects 2.7 TB of unusual outbound transfer from a data-processing VM over 46 minutes; flow logs identify 4 external IPs, disk snapshots show a new tool, and the VM handles 18% of nightly processing. How do you contain exfiltration and gate recovery?

    snapshotconcurrency
  • 80

    During an investigation, Azure Activity Log exports are missing for 3 subscriptions from 02:10 to 03:05, while ingestion metrics fell to 0 and 14 privileged changes occurred in that window. How do you preserve evidence, restore logging, and gate closure?

    loggingmonitoringclosures
  • 81

    A storage account containing 31 million customer documents allowed anonymous Blob access for 22 minutes; logs show 640,000 successful reads from 1,900 IPs and Defender reports mass download behavior. What do you contain, and what gates reopening?

  • 82

    AKS audit logs show a compromised pod creating cluster-admin bindings and reading 1,400 secrets in 6 minutes; Defender for Containers flags a privileged escape attempt on 3 nodes. How do you contain escalation and gate cluster recovery?

    containerskubernetessecrets
  • 83

    A Terraform state lock failed during 2 concurrent production applies; the remote state now lacks 86 resources that still exist, and both plans propose 240 replacements. How do you contain the state incident and gate the next apply?

    terraformincidentsconcurrency
  • 84

    An AzureRM provider upgrade changed defaults for 74 resources; a production plan shows 31 forced replacements and pipeline logs show the lock file was regenerated 18 minutes ago. Do you apply, pin, or edit state, and what gates recovery?

    packagingci-cd
  • 85

    A shared Terraform module release changed a subnet expression and 26 subscriptions now plan deletion of 190 subnets containing 4,800 private endpoints. What do you freeze, and how do you gate module recovery?

    networkingprivate-connectivityterraform
  • 86

    A Bicep deployment updated 420 resources across 12 resource groups before failing at resource 311; Activity Logs show 68 configuration changes and 9 applications now return 503. How do you roll back safely and prove recovery?

    deploymentconfigrollback
  • 87

    A DeployIfNotExists remediation across 540 subscriptions starts 20,000 ARM deployments; requests are being throttled and resources are appearing in workloads that should have been exempt. How do you stop and contain the rollout, reconcile its effects, and gate a restart?

    throttledeployment
  • 88

    An Azure DevOps workload identity was used from an unknown agent to deploy 17 resources and add Owner on 6 subscriptions in 11 minutes; pipeline audit logs show a token issued outside the approved pool. How do you contain the CI identity and gate pipeline recovery?

    tokensidentitydeployment
  • 89

    Azure Monitor missed a 19-minute checkout outage because Application Insights sampling dropped 62% of dependency spans, while 14,000 low-value alerts fired from a new threshold. How do you restore signal, contain noise, and gate monitoring recovery?

    monitoringdependenciesalerting
  • 90

    A canary deployment sent 5% of 50,000 requests per second to a new build, but its 8% error rate did not trigger rollback for 12 minutes because the alert evaluated the fleet-wide average. What do you roll back, and what gate must replace the failed canary check?

    deploymentrollbackalerting
  • 91

    West Europe has been unavailable for 13 minutes; the platform target is RTO 10 minutes and RPO 60 seconds, North Europe is at 45% warm capacity, and Front Door probes plus Service Health confirm regional failure. How do you activate DR and gate traffic?

    capacityedge-routing
  • 92

    Azure Backup has restored only 3.2 TB of a 14 TB SQL workload after 4 hours; RTO is 6 hours, projected completion is 11 hours, and a protected native restore point from 13 minutes before the incident is available. What do you do, and what gates service recovery?

    incidentssqlbackups
  • 93

    Azure Cost Management forecasts a 40% monthly increase from $500,000 to $700,000 after a release; exports show NAT Gateway and Log Analytics account for $165,000 of the variance, and finance needs containment in 6 hours. What do you cut, and what gates safety?

    gatewaynetworkingdispersion
  • 94

    A $240,000 monthly Azure reservation portfolio is only 54% utilized after 300 VMs moved size and region; the CFO review is tomorrow, and reverting workloads would add 80 ms latency. Do you move workloads, exchange reservations, or accept waste, and what gates the decision?

    latency
  • 95

    A regional vCPU quota blocks 420 VMSS instances during a 3-times traffic spike; allocation failures reach 100%, healthy capacity is 58%, and API errors are 17%. How do you contain the quota outage and gate added capacity?

    capacityapi
  • 96

    A landing-zone update changed 18 management-group policy assignments; within 2 hours, 96 subscriptions report broken deployments and 310 compliant private endpoints are denied. Policy Insights and Activity Logs identify the new initiative version. What do you roll back, and what gates restoration?

    deploymentrollbackendpoints
  • 97

    An audit begins in 72 hours, but 1,800 of 12,000 Azure resources lack required diagnostic settings and 340 storage accounts allow public network access; production teams report that bulk remediation could break 26 services. How do you meet the deadline without causing an outage?

    estimation
  • 98

    A severity-1 incident affects 42% of transactions across 3 regions; 28 responders are in one call, 11 changes were attempted in 20 minutes, and no one owns rollback while error rate rises to 19%. As incident commander, what do you do in the next 10 minutes?

    incidentstransactionsrollback
  • 99

    A junior engineer proposes granting Contributor to all 85 production subscriptions to fix a 25-minute deployment outage; 14 services are blocked, but audit logs show the actual denial comes from one custom role missing 2 actions. How do you mentor them while restoring service?

    mentoringdeployment
  • 100

    A 47-minute outage caused $1.8 million in failed orders; the postmortem lists 36 actions, but only 4 engineers are available this quarter and evidence shows 82% of impact came from a missing rollback alert and a shared regional DNS dependency. What do you prioritize?

    dependenciesrollbackalerting