Platform Engineer interview questions
100 real questions with model answers and explanations for Staff Platform Engineer candidates.
See a Platform Engineer resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would make a platform API the durable control plane and treat the portal as one replaceable client.
- Versioned service and environment APIs validate requests, persist desired state, and start asynchronous workflows with idempotency keys.
- Backstage exposes catalog, templates, status, and documentation, while Git and cloud controllers remain sources of reconciliation evidence.
- I would measure successful provisioning within 15 minutes, API availability, and rollback success per tenant rather than portal page uptime alone.
Why interviewers ask this: This tests whether the candidate separates an IDP product from its portal and designs around developer journeys.
I would keep every state-changing capability behind a documented platform API and leave presentation and discovery in the portal.
- The API owns authorization, validation, idempotency, audit records, and workflow state; Backstage plugins never call cloud APIs directly.
- The portal owns forms, catalog search, documentation, and status views, and CLI clients use the same API contracts.
- Contract tests and a 12-month deprecation policy let the UI change without breaking automation used by 80 teams.
Why interviewers ask this: A strong answer preserves stable automation contracts instead of coupling the platform to Backstage.
I would release the golden path as a versioned product with an automated upgrade path, not overwrite one template.
- New services receive v2, while v1 repositories get generated pull requests that show manifest, pipeline, and runtime changes.
- Compatibility tests run representative services before rollout, and 10 canaries must meet deploy and error-rate gates for 7 days.
- A dashboard tracks v1 owners, blocked migrations, and remaining days; exceptions have owners and expiry dates.
Why interviewers ask this: This checks product lifecycle ownership across templates and already-generated services.
I would provide a typed extension point with explicit ownership rather than permit arbitrary edits to generated infrastructure.
- Teams request a bounded network profile through the platform API, with policy checks and a documented support tier.
- Custom resources live in a separate tenant-owned layer so platform upgrades do not overwrite them.
- I would review the 36 exceptions quarterly; repeated needs graduate into the supported product, while stale exceptions expire.
Why interviewers ask this: This tests whether autonomy is bounded by clear security, support, and lifecycle rules.
I would run both contract versions long enough to migrate safely and make remaining risk visible by owner.
- Traffic telemetry identifies each client, request shape, and deprecated field before any deadline is announced.
- A compatibility adapter and generated migration guide cover the 60-day window, with office hours for the 19 blocked teams.
- Removal requires zero deprecated calls for 14 days; time-limited exceptions include a named owner and final date.
Why interviewers ask this: The interviewer is evaluating evidence-based deprecation rather than deadline-only governance.
I would first segment failed journeys, then remove the highest-friction boundary instead of adding more portal features.
- Funnel data compares catalog discovery, template start, first successful deploy, and repeat deploy by team and workload type.
- I would interview 8 adopting and 8 non-adopting teams, then fix one measured blocker such as a 45-minute approval wait.
- A 4-week pilot must improve time-to-first-deploy and repeat usage before the change becomes another golden-path option.
Why interviewers ask this: This tests whether the platform is managed as a measured engineer-facing product.
I would use distributed entity ownership in Git with centrally enforced schemas and incremental ingestion.
- Catalog processors validate system, owner, lifecycle, and dependency references before merging changes.
- Webhooks update changed locations, while a daily full reconciliation catches missed events without rescanning on every request.
- Freshness, orphan count, processor latency, and entities without active owners are platform SLIs exposed by domain.
Why interviewers ask this: A strong answer covers catalog data quality and reconciliation, not just portal installation.
I would isolate plugin dependencies and move slow integrations off the browser request path.
- Plugins consume stable backend APIs with timeouts, caching, and circuit breakers rather than querying 15 systems from the page.
- Bundle budgets and synthetic tests block a plugin that pushes p95 above 2 seconds on representative pages.
- Each plugin has an owner, support tier, permission scope, and removal path so unused integrations do not become permanent platform debt.
Why interviewers ask this: This checks performance isolation and lifecycle governance in a growing portal ecosystem.
I would model scaffolding as an idempotent workflow with durable step state and compensating cleanup.
- A request key binds repository, service identity, and environment so a retry resumes rather than creates a second resource.
- Each step records external identifiers before advancing, and reconciliation verifies Git, catalog, CI, and cloud state.
- Failed workflows show partial resources and safe retry or cleanup actions, with duplicate creation held below 0.1%.
Why interviewers ask this: This evaluates distributed workflow correctness across systems without a shared transaction.
I would use namespaces for ordinary trusted workloads and dedicated clusters where the 15-service blast-radius or security boundary requires it.
- Namespace tenants get RBAC, default-deny NetworkPolicy, ResourceQuota, Pod Security admission, and separate workload identities.
- Regulated, privileged, or high-noise workloads move to dedicated clusters because namespaces do not isolate the kernel or control plane.
- Placement policy records the reason, cost, and exit criteria so dedicated clusters remain an explicit tier rather than sprawl.
Why interviewers ask this: This tests whether multi-tenancy claims name actual isolation and blast-radius boundaries.
I would combine identity, network, admission, and resource boundaries because no single Kubernetes control provides tenant isolation.
- Namespace-scoped RBAC and workload identity prevent cross-tenant API and cloud access; default-deny policies restrict east-west traffic.
- ResourceQuota and LimitRange cap aggregate requests and limits, while priority classes protect platform services from tenant pressure.
- Admission rejects host access, privileged pods, and unapproved images, and boundary tests run on every fleet release.
Why interviewers ask this: The answer must distinguish security isolation from fair resource sharing.
I would treat clusters as replaceable fleet members driven by a versioned cluster API.
- A declared release channel pins Kubernetes, CNI, CSI, admission, and observability compatibility as one tested bundle.
- Upgrades move through 2 test, 3 canary, then regional waves, pausing on workload or platform SLO regression.
- Fleet inventory tracks version age, exceptions, and replacement readiness; clusters nearing 9 months cannot accept new tenants.
Why interviewers ask this: This checks fleet governance beyond running individual cluster upgrades.
I would adopt it only with a tested policy model and developer-visible flow evidence, not for eBPF alone.
- A canary cluster validates kube-proxy replacement, cloud networking, NetworkPolicy semantics, and upgrade compatibility under real traffic.
- Hubble flow views are scoped by tenant and linked from deployment diagnostics so teams can identify a denied path within 20 minutes.
- Rollout stops on DNS, connection, or policy-denial regressions, with the previous datapath retained until 10 clusters are stable.
Why interviewers ask this: This evaluates CNI adoption through developer impact and fleet safety.
I would offer the mesh only to journeys needing uniform mTLS or traffic policy and prove its cost against the 5 ms budget.
- A representative benchmark measures proxy latency, CPU, connection churn, and failure behavior before fleet rollout.
- Identity, retry, and timeout defaults are centrally versioned, while teams own application-level idempotency and latency budgets.
- Adoption begins with 20 services and expands only if p95 overhead and incident rate stay inside agreed gates.
Why interviewers ask this: A strong answer avoids making service mesh complexity a default without measured value.
I would expose typed platform resources and keep Terraform execution asynchronous behind the API.
- Requests validate policy and ownership, create immutable run records, and return status rather than hold an HTTP connection.
- State is sharded by tenant and resource domain, with one queued writer per state and short-lived cloud credentials.
- Plan, apply, drift, and rollback evidence is retained; p95 measures the complete developer journey, including queue time.
Why interviewers ask this: This tests safe Terraform orchestration rather than wrapping terraform apply in a web endpoint.
I would shard by ownership and lifecycle boundary, not place the fleet in one state or one state per tiny resource.
- Team-environment states cap lock contention and blast radius, while shared network and identity states have dedicated owners.
- Remote state uses versioning, encryption, locking, and audited recovery; plans are regenerated after any state or configuration change.
- Lock wait, state size, apply duration, and affected owners guide splitting before any shard reaches the 10-team limit.
Why interviewers ask this: This checks Terraform concurrency and recovery trade-offs with a defined ownership boundary.
I would choose Crossplane for continuously reconciled, reusable resource products where Kubernetes API semantics fit the ownership model.
- Compositions expose a small platform contract and hide provider resources, credentials, and policy defaults from tenants.
- Controller capacity, provider rate limits, and reconcile latency are load-tested against 10 minutes before broad adoption.
- Database changes needing careful one-time sequencing may remain in Terraform because permanent reconciliation can create ownership conflict.
Why interviewers ask this: The candidate should match the control model to the resource lifecycle rather than choose by tool preference.
I would establish one authoritative writer per resource before enabling reconciliation.
- Inventory maps resource IDs, state addresses, Crossplane references, and team owners; ambiguous resources are frozen from change.
- Migration imports one bounded group, verifies no diff, then removes it from the old controller before Crossplane management starts.
- Admission policy rejects resources with the old ownership label, and 14 days of zero conflicting writes closes each wave.
Why interviewers ask this: This tests safe controller handoff without dual ownership.
I would support Pulumi only through the same small infrastructure contracts, not as a second unrestricted provisioning stack.
- Teams may author approved components in TypeScript, but identity, policy, state, and audit requirements remain platform-owned.
- A 3-team pilot measures support tickets, change lead time, and failed applies against Terraform equivalents for 8 weeks.
- If it doubles operational paths without a measured developer gain, I would keep Terraform as the supported implementation.
Why interviewers ask this: This evaluates bounded optionality under a concrete support constraint.
I would abstract only the shared application contract and keep provider-specific capabilities explicit.
- The platform API offers compute, identity, observability, and data classes with region and residency policy encoded in admission.
- EU workloads can fail over only between the 2 approved regions; replication and recovery tests prove the 30-minute RTO.
- Provider-specific extensions are visible escape hatches so lowest-common-denominator abstractions do not hide reliability differences.
Why interviewers ask this: Multi-cloud is justified here by named residency and recovery constraints rather than generic portability.
Locked questions
- 21
Design GitOps for 50 clusters, 800 applications, and a requirement that a bad change reaches at most 2 clusters.
gitopsdesigncluster - 22
How would you progressively roll out a platform admission change to 45 clusters with a 0.5% rejection-error budget?
reliabilityclustererror-budget - 23
How would you partition Argo CD for 1,000 apps across 40 clusters while keeping tenant visibility isolated?
gitopspartitioningcss - 24
Design CI/CD as a shared product for 2,500 engineers, 30,000 jobs per day, and 99.9% job-start availability.
designavailabilityci-cd - 25
Choose runner isolation for 150 teams and 20 untrusted repositories when startup p95 must stay below 45 seconds.
- 26
How would you size CI capacity for a 10x morning burst from 200 to 2,000 queued jobs with p95 wait under 3 minutes?
capacitydata-structurescapacity-planning - 27
Design artifact storage for 25 TB per month, 90-day retention, and restores under 10 minutes.
retentiondesignartifacts - 28
How would you enforce provenance, SBOM, and signing across 800 services with fewer than 1% blocked good releases?
supply-chain - 29
Design policy-as-code feedback for 500 developers when 95% of violations must be understood without a support ticket.
feedbackdesign - 30
Where would you enforce 35 infrastructure policies across CI, Terraform, and Kubernetes while adding under 30 seconds to feedback?
kubernetesterraformfeedback - 31
How would you govern 70 policy exceptions when each may last no more than 30 days?
error-handling - 32
Design workload identity for 600 services across 25 clusters with credentials valid for at most 1 hour.
designcluster - 33
How would you rotate 4,000 secrets every 30 days without restarting all 700 workloads at once?
secretsconfiguration - 34
Design observability as a platform for 200 teams producing 8 million samples per second with tenant query isolation.
designobservabilityqueries - 35
Prometheus cardinality must stay below 20 million active series across 500 services; what platform controls do you design?
monitoringdesign - 36
How would you keep log ingestion under 120 TB per month for 600 services without hiding production failures?
- 37
Design trace sampling for 1 million requests per second while retaining 99% of traces for failed critical journeys.
samplingdesign - 38
Define platform SLIs for 1,500 developers when the target is 99.9% successful first deploy within 15 minutes.
deployment - 39
A 99.9% deployment-journey SLO has consumed 60% of its 30-day error budget in 5 days; what release policy do you design?
designreliabilityslo - 40
How would you attribute a 12-minute provisioning SLO when cloud APIs account for 8 minutes of p95 latency?
slolatencyapi - 41
How would you reduce median time-to-first-deploy from 3 days to 4 hours for 25 new teams per quarter?
deployment - 42
How would you measure cognitive load for 900 developers when a service requires editing 14 configuration files?
config - 43
A 7-person platform team handles 320 support tickets per month; how do you cut toil by 50% in 2 quarters?
reliabilityplatform-engineeringtoil - 44
Beyond the 5 modern DORA metrics, including deployment rework rate, what would you track for a platform serving 60 teams and 2,000 weekly deployments?
deploymentmonitoringworkloads - 45
How would you test whether a new golden path improves adoption from 45% to 70% across 40 teams in 90 days?
golden-pathdecision-making - 46
Design showback for 150 teams and $900,000 monthly cloud spend when shared costs are 22% of the bill.
design - 47
How would you improve resource efficiency by 25% across 500 services without increasing p99 latency above 200 ms?
latency - 48
Design disaster recovery for the platform API and GitOps control plane with a 30-minute RTO and 5-minute RPO across 2 regions.
designdisaster-recoverycontrol-plane - 49
How would you preserve compatibility for 350 services during a 6-month platform API and Kubernetes CRD migration?
kubernetesapimigrations - 50
For 2,500 engineers and 1,000 services, which 3 platform boundaries would you standardize first if support capacity is 8 engineers?
capacitycapacity-planning - 51
Backstage p95 latency rose from 1.4 to 9 seconds for 1,600 users after plugin release 32; what do you do?
latencybackstage - 52
Catalog ownership is stale for 280 of 12,000 entities and 45 alerts routed to departed teams in 7 days; how do you recover?
ownershipalerting - 53
Backstage search stopped returning 68% of 12,000 catalog entities after indexer release 41, blocking 310 developers from finding services and runbooks in 2 hours; how do you recover?
runbooksbackstageindexes - 54
After queue schema v8, 63 asynchronous completion callbacks were lost and 41 developers saw requests stuck although all 41 resources existed; how do you recover?
asynccallbacksdata-structures - 55
The platform API returns 503 for 38% of 900 provisioning requests during 22 minutes, but Backstage remains green; how do you lead the incident?
incidentsincident-managementapi - 56
Golden-path v4 raised failed deployments from 1.2% to 17% across 74 of 310 services in 35 minutes; what do you decide?
kubernetes-deploymentdeploymentworkloads - 57
A Backstage plugin exposed metadata from 12 tenants to 3 unauthorized teams for 46 minutes; what evidence and controls do you require?
backstage - 58
Reusable CI workflow v9 broke builds in 186 of 740 repositories within 28 minutes; how do you restore developer releases and change the rollout?
- 59
A shared runner executed unknown code for 11 minutes and could read caches from 85 repositories; what is your containment plan?
caching - 60
Checksum failures affect 37 of 4,200 artifacts replicated to region 2, and 9 deployments consumed them; what do you do?
deploymentartifactsreplication - 61
A Terraform runner selected the wrong workspace and cloud account, applying 37 production changes from 1 staging platform request; how do you contain and recover?
terraform - 62
A Terraform state or plan artifact exposed 26 plaintext secrets to 14 engineers for 52 minutes; how do you contain it and prevent another platform leak?
terraformartifactssecrets - 63
Terraform provider 6.2 proposes replacement of 140 databases where 6.1 showed no change; what rollout decision do you make?
terraformdatabase - 64
Crossplane reconcile traffic rose from 80 to 9,000 requests per minute and hit cloud API 429s for 23 minutes; how do you stabilize it?
iacapi - 65
Crossplane and Terraform alternated tags on 96 resources every 3 minutes, creating 1,800 audit events; how do you end the conflict?
terraformiac - 66
Argo CD reports 210 apps OutOfSync, but live manifests match Git for 198; how do you investigate without mass sync?
gitgitops - 67
An ApplicationSet change triggered 4,800 syncs across 40 clusters in 6 minutes and saturated 9 API servers; what do you do?
control-planeclusterapi - 68
A bad promotion reached 7 of 50 clusters despite a 2-cluster limit and caused 61 failed apps; how do you repair the rollout design?
designcluster - 69
Kubernetes API p99 rose from 180 ms to 4.8 seconds in 3 shared clusters, causing 32% platform request failures; what do you check?
kubernetesclusterapi - 70
DNS failures rose from 0.02% to 8% across 420 services after a fleet change, while 6 canary clusters looked healthy; how do you respond?
dnsdeployment-strategiescluster - 71
Cilium IPAM exhausted address pools in 8 clusters, leaving 620 developer deployment pods unschedulable for 27 minutes; how do you restore the platform journey?
workloadsclusterdeployment - 72
A CRD conversion webhook times out for 14% of 60,000 objects during an API migration; how do you avoid data loss?
webhooksmigrations - 73
Kubernetes upgrade 1.35 increased pod startup p95 from 22 to 95 seconds in 4 of 30 clusters; what rollout decision do you make?
workloadskubernetescluster - 74
One tenant consumed 62% actual CPU and pushed p99 above 1 second for 47 services in a 90-tenant cluster; how do you contain it?
cluster - 75
Certificates for 130 services expire within 9 hours because rotation stalled 2 days ago; what do you do?
- 76
A NetworkPolicy rollout blocked payments from 23 services for 11 minutes and Hubble shows 48,000 denied flows; how do you recover?
network-policy - 77
A 1-zone outage took 36 platform-managed services offline and blocked 240 developer deployments despite 3 replicas each; how do you fix the topology defaults?
kubernetes-deploymentreplicationdeployment - 78
HPA scaled 70 workloads from 2 to 10 replicas, but 430 pods stayed Pending for 16 minutes because cluster capacity was unschedulable; how do you restore fleet capacity?
capacityclusterreplication - 79
A service-mesh retry policy tripled traffic to 1.8 million requests per second and raised errors from 2% to 31%; what do you do?
resilience - 80
mTLS failures reached 12% for 260 services after trust bundle v17, but 18 services cannot roll back; how do you recover?
mtlsrollback - 81
A secret used by 74 workloads appeared in CI logs for 26 minutes and was viewed 19 times; what is your response?
secretsconfiguration - 82
An audit finds 3,800 Kubernetes Secrets stored only as base64 across 25 clusters; what migration do you lead in 60 days?
configurationkubernetescluster - 83
A forged image signature reached 6 production clusters and ran for 17 minutes in 14 pods; how do you contain the supply-chain incident?
incidentsincident-managementcluster - 84
OpenTelemetry collectors dropped 28% of spans for 95 services during a 40-minute burst; how do you restore trustworthy telemetry?
observability - 85
A remote-write gateway acknowledged 2.4 billion samples but lost 18% for 46 tenants during a 32-minute metrics-backend outage; how do you repair trust?
gatewaymonitoring - 86
Log cost rose from $180,000 to $510,000 in 1 month while query evidence shows 67% of new data was never read; what do you change?
queries - 87
The global platform SLO remains green while 38% of first-deploy attempts fail for 42 EU teams over 6 hours; what incident and measurement changes do you make?
incidentsincident-managementdeployment - 88
A cloud identity API caused 72% of 1,100 provisioning failures for 31 minutes; how do you attribute and mitigate the dependency safely?
dependenciesapi - 89
Moving CI and developer-preview pools to 80% spot capacity saved $210,000 in 1 month but interruptions canceled 1,900 jobs and 340 previews; what do you change?
capacitycapacity-planning - 90
A team disputes $84,000 of a $310,000 monthly showback because 27% is shared observability cost; how do you resolve it?
observability - 91
A regional outage lasts 47 minutes, exceeding the platform RTO of 30 minutes, and 180 deployments are pending; how do you recover?
disaster-recoveryworkloadsdeployment - 92
A restore meets the 5-minute RPO but loses 37 workflow status records and shows 12 duplicate cloud accounts; what changes?
disaster-recovery - 93
TechDocs showed a database-recovery runbook 4 versions behind production, causing 26 teams to fail 41 self-service drills in 6 days; how do you prevent recurrence?
runbooksdocumentationdatabase - 94
Platform scorecard rule v12 marked 73 noncompliant services compliant and blocked releases for 28 compliant services in 3 hours; how do you repair decisions?
- 95
Rollback failed for 22 services because artifact retention deleted production images after 30 days while release records remained for 90 days; how do you recover?
rollbackartifactsretention - 96
Why are 37 expired policy exemptions still active after agents cached an old OPA bundle for 19 hours across 16 clusters, and how do you contain the platform risk?
policycachingcluster - 97
A platform SDK v3 migration breaks 6% of 420 services because clients relied on an undocumented timeout; how do you restore compatibility?
migrationsresilience - 98
A mentee submits a Backstage plugin that lets 1 ordinary user edit all 600 catalog entities; how do you review it and prove resource-level authorization in 4 weeks?
authbackstage - 99
A senior teammate's incident review lists 12 actions but no evidence, after a 43-minute CI outage affecting 2,300 jobs; how do you mentor them?
mentoringincidentsincident-management - 100
A mentee's golden-path change cut setup from 6 hours to 50 minutes but raised failed first deploys from 3% to 11%; how do you guide the decision?
deployment