Skip to content

Platform Engineer interview questions

100 real questions with model answers and explanations for Staff Platform Engineer candidates.

See a Platform Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

designslocontrol-plane

I would make a platform API the durable control plane and treat the portal as one replaceable client.

  • Versioned service and environment APIs validate requests, persist desired state, and start asynchronous workflows with idempotency keys.
  • Backstage exposes catalog, templates, status, and documentation, while Git and cloud controllers remain sources of reconciliation evidence.
  • I would measure successful provisioning within 15 minutes, API availability, and rollback success per tenant rather than portal page uptime alone.

Why interviewers ask this: This tests whether the candidate separates an IDP product from its portal and designs around developer journeys.

api

I would keep every state-changing capability behind a documented platform API and leave presentation and discovery in the portal.

  • The API owns authorization, validation, idempotency, audit records, and workflow state; Backstage plugins never call cloud APIs directly.
  • The portal owns forms, catalog search, documentation, and status views, and CLI clients use the same API contracts.
  • Contract tests and a 12-month deprecation policy let the UI change without breaking automation used by 80 teams.

Why interviewers ask this: A strong answer preserves stable automation contracts instead of coupling the platform to Backstage.

golden-path

I would release the golden path as a versioned product with an automated upgrade path, not overwrite one template.

  • New services receive v2, while v1 repositories get generated pull requests that show manifest, pipeline, and runtime changes.
  • Compatibility tests run representative services before rollout, and 10 canaries must meet deploy and error-rate gates for 7 days.
  • A dashboard tracks v1 owners, blocked migrations, and remaining days; exceptions have owners and expiry dates.

Why interviewers ask this: This checks product lifecycle ownership across templates and already-generated services.

golden-pathdesignescape-hatches

I would provide a typed extension point with explicit ownership rather than permit arbitrary edits to generated infrastructure.

  • Teams request a bounded network profile through the platform API, with policy checks and a documented support tier.
  • Custom resources live in a separate tenant-owned layer so platform upgrades do not overwrite them.
  • I would review the 36 exceptions quarterly; repeated needs graduate into the supported product, while stale exceptions expire.

Why interviewers ask this: This tests whether autonomy is bounded by clear security, support, and lifecycle rules.

api

I would run both contract versions long enough to migrate safely and make remaining risk visible by owner.

  • Traffic telemetry identifies each client, request shape, and deprecated field before any deadline is announced.
  • A compatibility adapter and generated migration guide cover the 60-day window, with office hours for the 19 blocked teams.
  • Removal requires zero deprecated calls for 14 days; time-limited exceptions include a named owner and final date.

Why interviewers ask this: The interviewer is evaluating evidence-based deprecation rather than deadline-only governance.

self-servicearchitecturedecision-making

I would first segment failed journeys, then remove the highest-friction boundary instead of adding more portal features.

  • Funnel data compares catalog discovery, template start, first successful deploy, and repeat deploy by team and workload type.
  • I would interview 8 adopting and 8 non-adopting teams, then fix one measured blocker such as a 45-minute approval wait.
  • A 4-week pilot must improve time-to-first-deploy and repeat usage before the change becomes another golden-path option.

Why interviewers ask this: This tests whether the platform is managed as a measured engineer-facing product.

ownershipdesignbackstage

I would use distributed entity ownership in Git with centrally enforced schemas and incremental ingestion.

  • Catalog processors validate system, owner, lifecycle, and dependency references before merging changes.
  • Webhooks update changed locations, while a daily full reconciliation catches missed events without rescanning on every request.
  • Freshness, orphan count, processor latency, and entities without active owners are platform SLIs exposed by domain.

Why interviewers ask this: A strong answer covers catalog data quality and reconciliation, not just portal installation.

backstage

I would isolate plugin dependencies and move slow integrations off the browser request path.

  • Plugins consume stable backend APIs with timeouts, caching, and circuit breakers rather than querying 15 systems from the page.
  • Bundle budgets and synthetic tests block a plugin that pushes p95 above 2 seconds on representative pages.
  • Each plugin has an owner, support tier, permission scope, and removal path so unused integrations do not become permanent platform debt.

Why interviewers ask this: This checks performance isolation and lifecycle governance in a growing portal ecosystem.

I would model scaffolding as an idempotent workflow with durable step state and compensating cleanup.

  • A request key binds repository, service identity, and environment so a retry resumes rather than creates a second resource.
  • Each step records external identifiers before advancing, and reconciliation verifies Git, catalog, CI, and cloud state.
  • Failed workflows show partial resources and safe retry or cleanup actions, with duplicate creation held below 0.1%.

Why interviewers ask this: This evaluates distributed workflow correctness across systems without a shared transaction.

resiliencenamespacescluster

I would use namespaces for ordinary trusted workloads and dedicated clusters where the 15-service blast-radius or security boundary requires it.

  • Namespace tenants get RBAC, default-deny NetworkPolicy, ResourceQuota, Pod Security admission, and separate workload identities.
  • Regulated, privileged, or high-noise workloads move to dedicated clusters because namespaces do not isolate the kernel or control plane.
  • Placement policy records the reason, cost, and exit criteria so dedicated clusters remain an explicit tier rather than sprawl.

Why interviewers ask this: This tests whether multi-tenancy claims name actual isolation and blast-radius boundaries.

designkubernetescluster

I would combine identity, network, admission, and resource boundaries because no single Kubernetes control provides tenant isolation.

  • Namespace-scoped RBAC and workload identity prevent cross-tenant API and cloud access; default-deny policies restrict east-west traffic.
  • ResourceQuota and LimitRange cap aggregate requests and limits, while priority classes protect platform services from tenant pressure.
  • Admission rejects host access, privileged pods, and unapproved images, and boundary tests run on every fleet release.

Why interviewers ask this: The answer must distinguish security isolation from fair resource sharing.

kubernetescluster

I would treat clusters as replaceable fleet members driven by a versioned cluster API.

  • A declared release channel pins Kubernetes, CNI, CSI, admission, and observability compatibility as one tested bundle.
  • Upgrades move through 2 test, 3 canary, then regional waves, pausing on workload or platform SLO regression.
  • Fleet inventory tracks version age, exceptions, and replacement readiness; clusters nearing 9 months cannot accept new tenants.

Why interviewers ask this: This checks fleet governance beyond running individual cluster upgrades.

decision-makingcluster

I would adopt it only with a tested policy model and developer-visible flow evidence, not for eBPF alone.

  • A canary cluster validates kube-proxy replacement, cloud networking, NetworkPolicy semantics, and upgrade compatibility under real traffic.
  • Hubble flow views are scoped by tenant and linked from deployment diagnostics so teams can identify a denied path within 20 minutes.
  • Rollout stops on DNS, connection, or policy-denial regressions, with the previous datapath retained until 10 clusters are stable.

Why interviewers ask this: This evaluates CNI adoption through developer impact and fleet safety.

service-meshlatency

I would offer the mesh only to journeys needing uniform mTLS or traffic policy and prove its cost against the 5 ms budget.

  • A representative benchmark measures proxy latency, CPU, connection churn, and failure behavior before fleet rollout.
  • Identity, retry, and timeout defaults are centrally versioned, while teams own application-level idempotency and latency budgets.
  • Adoption begins with 20 services and expands only if p95 overhead and incident rate stay inside agreed gates.

Why interviewers ask this: A strong answer avoids making service mesh complexity a default without measured value.

terraformapidesign

I would expose typed platform resources and keep Terraform execution asynchronous behind the API.

  • Requests validate policy and ownership, create immutable run records, and return status rather than hold an HTTP connection.
  • State is sharded by tenant and resource domain, with one queued writer per state and short-lived cloud credentials.
  • Plan, apply, drift, and rollback evidence is retained; p95 measures the complete developer journey, including queue time.

Why interviewers ask this: This tests safe Terraform orchestration rather than wrapping terraform apply in a web endpoint.

terraformsharding

I would shard by ownership and lifecycle boundary, not place the fleet in one state or one state per tiny resource.

  • Team-environment states cap lock contention and blast radius, while shared network and identity states have dedicated owners.
  • Remote state uses versioning, encryption, locking, and audited recovery; plans are regenerated after any state or configuration change.
  • Lock wait, state size, apply duration, and affected owners guide splitting before any shard reaches the 10-team limit.

Why interviewers ask this: This checks Terraform concurrency and recovery trade-offs with a defined ownership boundary.

iac

I would choose Crossplane for continuously reconciled, reusable resource products where Kubernetes API semantics fit the ownership model.

  • Compositions expose a small platform contract and hide provider resources, credentials, and policy defaults from tenants.
  • Controller capacity, provider rate limits, and reconcile latency are load-tested against 10 minutes before broad adoption.
  • Database changes needing careful one-time sequencing may remain in Terraform because permanent reconciliation can create ownership conflict.

Why interviewers ask this: The candidate should match the control model to the resource lifecycle rather than choose by tool preference.

terraformiacownership

I would establish one authoritative writer per resource before enabling reconciliation.

  • Inventory maps resource IDs, state addresses, Crossplane references, and team owners; ambiguous resources are frozen from change.
  • Migration imports one bounded group, verifies no diff, then removes it from the old controller before Crossplane management starts.
  • Admission policy rejects resources with the old ownership label, and 14 days of zero conflicting writes closes each wave.

Why interviewers ask this: This tests safe controller handoff without dual ownership.

terraformiaccapacity

I would support Pulumi only through the same small infrastructure contracts, not as a second unrestricted provisioning stack.

  • Teams may author approved components in TypeScript, but identity, policy, state, and audit requirements remain platform-owned.
  • A 3-team pilot measures support tickets, change lead time, and failed applies against Terraform equivalents for 8 weeks.
  • If it doubles operational paths without a measured developer gain, I would keep Terraform as the supported implementation.

Why interviewers ask this: This evaluates bounded optionality under a concrete support constraint.

designfailoverdisaster-recovery

I would abstract only the shared application contract and keep provider-specific capabilities explicit.

  • The platform API offers compute, identity, observability, and data classes with region and residency policy encoded in admission.
  • EU workloads can fail over only between the 2 approved regions; replication and recovery tests prove the 30-minute RTO.
  • Provider-specific extensions are visible escape hatches so lowest-common-denominator abstractions do not hide reliability differences.

Why interviewers ask this: Multi-cloud is justified here by named residency and recovery constraints rather than generic portability.

Locked questions

  • 21

    Design GitOps for 50 clusters, 800 applications, and a requirement that a bad change reaches at most 2 clusters.

    gitopsdesigncluster
  • 22

    How would you progressively roll out a platform admission change to 45 clusters with a 0.5% rejection-error budget?

    reliabilityclustererror-budget
  • 23

    How would you partition Argo CD for 1,000 apps across 40 clusters while keeping tenant visibility isolated?

    gitopspartitioningcss
  • 24

    Design CI/CD as a shared product for 2,500 engineers, 30,000 jobs per day, and 99.9% job-start availability.

    designavailabilityci-cd
  • 25

    Choose runner isolation for 150 teams and 20 untrusted repositories when startup p95 must stay below 45 seconds.

  • 26

    How would you size CI capacity for a 10x morning burst from 200 to 2,000 queued jobs with p95 wait under 3 minutes?

    capacitydata-structurescapacity-planning
  • 27

    Design artifact storage for 25 TB per month, 90-day retention, and restores under 10 minutes.

    retentiondesignartifacts
  • 28

    How would you enforce provenance, SBOM, and signing across 800 services with fewer than 1% blocked good releases?

    supply-chain
  • 29

    Design policy-as-code feedback for 500 developers when 95% of violations must be understood without a support ticket.

    feedbackdesign
  • 30

    Where would you enforce 35 infrastructure policies across CI, Terraform, and Kubernetes while adding under 30 seconds to feedback?

    kubernetesterraformfeedback
  • 31

    How would you govern 70 policy exceptions when each may last no more than 30 days?

    error-handling
  • 32

    Design workload identity for 600 services across 25 clusters with credentials valid for at most 1 hour.

    designcluster
  • 33

    How would you rotate 4,000 secrets every 30 days without restarting all 700 workloads at once?

    secretsconfiguration
  • 34

    Design observability as a platform for 200 teams producing 8 million samples per second with tenant query isolation.

    designobservabilityqueries
  • 35

    Prometheus cardinality must stay below 20 million active series across 500 services; what platform controls do you design?

    monitoringdesign
  • 36

    How would you keep log ingestion under 120 TB per month for 600 services without hiding production failures?

  • 37

    Design trace sampling for 1 million requests per second while retaining 99% of traces for failed critical journeys.

    samplingdesign
  • 38

    Define platform SLIs for 1,500 developers when the target is 99.9% successful first deploy within 15 minutes.

    deployment
  • 39

    A 99.9% deployment-journey SLO has consumed 60% of its 30-day error budget in 5 days; what release policy do you design?

    designreliabilityslo
  • 40

    How would you attribute a 12-minute provisioning SLO when cloud APIs account for 8 minutes of p95 latency?

    slolatencyapi
  • 41

    How would you reduce median time-to-first-deploy from 3 days to 4 hours for 25 new teams per quarter?

    deployment
  • 42

    How would you measure cognitive load for 900 developers when a service requires editing 14 configuration files?

    config
  • 43

    A 7-person platform team handles 320 support tickets per month; how do you cut toil by 50% in 2 quarters?

    reliabilityplatform-engineeringtoil
  • 44

    Beyond the 5 modern DORA metrics, including deployment rework rate, what would you track for a platform serving 60 teams and 2,000 weekly deployments?

    deploymentmonitoringworkloads
  • 45

    How would you test whether a new golden path improves adoption from 45% to 70% across 40 teams in 90 days?

    golden-pathdecision-making
  • 46

    Design showback for 150 teams and $900,000 monthly cloud spend when shared costs are 22% of the bill.

    design
  • 47

    How would you improve resource efficiency by 25% across 500 services without increasing p99 latency above 200 ms?

    latency
  • 48

    Design disaster recovery for the platform API and GitOps control plane with a 30-minute RTO and 5-minute RPO across 2 regions.

    designdisaster-recoverycontrol-plane
  • 49

    How would you preserve compatibility for 350 services during a 6-month platform API and Kubernetes CRD migration?

    kubernetesapimigrations
  • 50

    For 2,500 engineers and 1,000 services, which 3 platform boundaries would you standardize first if support capacity is 8 engineers?

    capacitycapacity-planning
  • 51

    Backstage p95 latency rose from 1.4 to 9 seconds for 1,600 users after plugin release 32; what do you do?

    latencybackstage
  • 52

    Catalog ownership is stale for 280 of 12,000 entities and 45 alerts routed to departed teams in 7 days; how do you recover?

    ownershipalerting
  • 53

    Backstage search stopped returning 68% of 12,000 catalog entities after indexer release 41, blocking 310 developers from finding services and runbooks in 2 hours; how do you recover?

    runbooksbackstageindexes
  • 54

    After queue schema v8, 63 asynchronous completion callbacks were lost and 41 developers saw requests stuck although all 41 resources existed; how do you recover?

    asynccallbacksdata-structures
  • 55

    The platform API returns 503 for 38% of 900 provisioning requests during 22 minutes, but Backstage remains green; how do you lead the incident?

    incidentsincident-managementapi
  • 56

    Golden-path v4 raised failed deployments from 1.2% to 17% across 74 of 310 services in 35 minutes; what do you decide?

    kubernetes-deploymentdeploymentworkloads
  • 57

    A Backstage plugin exposed metadata from 12 tenants to 3 unauthorized teams for 46 minutes; what evidence and controls do you require?

    backstage
  • 58

    Reusable CI workflow v9 broke builds in 186 of 740 repositories within 28 minutes; how do you restore developer releases and change the rollout?

  • 59

    A shared runner executed unknown code for 11 minutes and could read caches from 85 repositories; what is your containment plan?

    caching
  • 60

    Checksum failures affect 37 of 4,200 artifacts replicated to region 2, and 9 deployments consumed them; what do you do?

    deploymentartifactsreplication
  • 61

    A Terraform runner selected the wrong workspace and cloud account, applying 37 production changes from 1 staging platform request; how do you contain and recover?

    terraform
  • 62

    A Terraform state or plan artifact exposed 26 plaintext secrets to 14 engineers for 52 minutes; how do you contain it and prevent another platform leak?

    terraformartifactssecrets
  • 63

    Terraform provider 6.2 proposes replacement of 140 databases where 6.1 showed no change; what rollout decision do you make?

    terraformdatabase
  • 64

    Crossplane reconcile traffic rose from 80 to 9,000 requests per minute and hit cloud API 429s for 23 minutes; how do you stabilize it?

    iacapi
  • 65

    Crossplane and Terraform alternated tags on 96 resources every 3 minutes, creating 1,800 audit events; how do you end the conflict?

    terraformiac
  • 66

    Argo CD reports 210 apps OutOfSync, but live manifests match Git for 198; how do you investigate without mass sync?

    gitgitops
  • 67

    An ApplicationSet change triggered 4,800 syncs across 40 clusters in 6 minutes and saturated 9 API servers; what do you do?

    control-planeclusterapi
  • 68

    A bad promotion reached 7 of 50 clusters despite a 2-cluster limit and caused 61 failed apps; how do you repair the rollout design?

    designcluster
  • 69

    Kubernetes API p99 rose from 180 ms to 4.8 seconds in 3 shared clusters, causing 32% platform request failures; what do you check?

    kubernetesclusterapi
  • 70

    DNS failures rose from 0.02% to 8% across 420 services after a fleet change, while 6 canary clusters looked healthy; how do you respond?

    dnsdeployment-strategiescluster
  • 71

    Cilium IPAM exhausted address pools in 8 clusters, leaving 620 developer deployment pods unschedulable for 27 minutes; how do you restore the platform journey?

    workloadsclusterdeployment
  • 72

    A CRD conversion webhook times out for 14% of 60,000 objects during an API migration; how do you avoid data loss?

    webhooksmigrations
  • 73

    Kubernetes upgrade 1.35 increased pod startup p95 from 22 to 95 seconds in 4 of 30 clusters; what rollout decision do you make?

    workloadskubernetescluster
  • 74

    One tenant consumed 62% actual CPU and pushed p99 above 1 second for 47 services in a 90-tenant cluster; how do you contain it?

    cluster
  • 75

    Certificates for 130 services expire within 9 hours because rotation stalled 2 days ago; what do you do?

  • 76

    A NetworkPolicy rollout blocked payments from 23 services for 11 minutes and Hubble shows 48,000 denied flows; how do you recover?

    network-policy
  • 77

    A 1-zone outage took 36 platform-managed services offline and blocked 240 developer deployments despite 3 replicas each; how do you fix the topology defaults?

    kubernetes-deploymentreplicationdeployment
  • 78

    HPA scaled 70 workloads from 2 to 10 replicas, but 430 pods stayed Pending for 16 minutes because cluster capacity was unschedulable; how do you restore fleet capacity?

    capacityclusterreplication
  • 79

    A service-mesh retry policy tripled traffic to 1.8 million requests per second and raised errors from 2% to 31%; what do you do?

    resilience
  • 80

    mTLS failures reached 12% for 260 services after trust bundle v17, but 18 services cannot roll back; how do you recover?

    mtlsrollback
  • 81

    A secret used by 74 workloads appeared in CI logs for 26 minutes and was viewed 19 times; what is your response?

    secretsconfiguration
  • 82

    An audit finds 3,800 Kubernetes Secrets stored only as base64 across 25 clusters; what migration do you lead in 60 days?

    configurationkubernetescluster
  • 83

    A forged image signature reached 6 production clusters and ran for 17 minutes in 14 pods; how do you contain the supply-chain incident?

    incidentsincident-managementcluster
  • 84

    OpenTelemetry collectors dropped 28% of spans for 95 services during a 40-minute burst; how do you restore trustworthy telemetry?

    observability
  • 85

    A remote-write gateway acknowledged 2.4 billion samples but lost 18% for 46 tenants during a 32-minute metrics-backend outage; how do you repair trust?

    gatewaymonitoring
  • 86

    Log cost rose from $180,000 to $510,000 in 1 month while query evidence shows 67% of new data was never read; what do you change?

    queries
  • 87

    The global platform SLO remains green while 38% of first-deploy attempts fail for 42 EU teams over 6 hours; what incident and measurement changes do you make?

    incidentsincident-managementdeployment
  • 88

    A cloud identity API caused 72% of 1,100 provisioning failures for 31 minutes; how do you attribute and mitigate the dependency safely?

    dependenciesapi
  • 89

    Moving CI and developer-preview pools to 80% spot capacity saved $210,000 in 1 month but interruptions canceled 1,900 jobs and 340 previews; what do you change?

    capacitycapacity-planning
  • 90

    A team disputes $84,000 of a $310,000 monthly showback because 27% is shared observability cost; how do you resolve it?

    observability
  • 91

    A regional outage lasts 47 minutes, exceeding the platform RTO of 30 minutes, and 180 deployments are pending; how do you recover?

    disaster-recoveryworkloadsdeployment
  • 92

    A restore meets the 5-minute RPO but loses 37 workflow status records and shows 12 duplicate cloud accounts; what changes?

    disaster-recovery
  • 93

    TechDocs showed a database-recovery runbook 4 versions behind production, causing 26 teams to fail 41 self-service drills in 6 days; how do you prevent recurrence?

    runbooksdocumentationdatabase
  • 94

    Platform scorecard rule v12 marked 73 noncompliant services compliant and blocked releases for 28 compliant services in 3 hours; how do you repair decisions?

  • 95

    Rollback failed for 22 services because artifact retention deleted production images after 30 days while release records remained for 90 days; how do you recover?

    rollbackartifactsretention
  • 96

    Why are 37 expired policy exemptions still active after agents cached an old OPA bundle for 19 hours across 16 clusters, and how do you contain the platform risk?

    policycachingcluster
  • 97

    A platform SDK v3 migration breaks 6% of 420 services because clients relied on an undocumented timeout; how do you restore compatibility?

    migrationsresilience
  • 98

    A mentee submits a Backstage plugin that lets 1 ordinary user edit all 600 catalog entities; how do you review it and prove resource-level authorization in 4 weeks?

    authbackstage
  • 99

    A senior teammate's incident review lists 12 actions but no evidence, after a 43-minute CI outage affecting 2,300 jobs; how do you mentor them?

    mentoringincidentsincident-management
  • 100

    A mentee's golden-path change cut setup from 6 hours to 50 minutes but raised failed first deploys from 3% to 11%; how do you guide the decision?

    deployment