Technical Program Manager interview questions
100 real questions with model answers and explanations for Senior candidates.
See a Technical Program Manager resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would treat both chains as schedule-controlling until actual delivery creates meaningful float on 1 of them.
- Within 48 hours, I build a directed graph with each deliverable, owner, duration range, acceptance condition, and latest safe date, then remove cycles before baselining it.
- I reserve engineering and review capacity for both chains through week 6, but expedite only work that advances an end-to-end checkout slice.
- Gates in weeks 4 and 7 recalculate both paths from completed work and test evidence; less than 5 working days of float triggers a scope decision.
- If both paths remain critical at week 7, I remove a 2-week loyalty edge case rather than add people late and increase integration load.
Why interviewers ask this: The interviewer is evaluating whether you can manage a changing critical path and make a quantified scope choice across 2 near-critical chains.
I would launch v2 in week 9 and fund a bounded dual-support period because the 2 slow consumers cannot safely meet that cutover.
- By day 5, the teams freeze an OpenAPI contract covering field semantics, errors, idempotency, rate limits, and 1 owner for every disputed behavior.
- Gates in weeks 3 and 6 require all 4 consumers to pass their declared version contract against the provider build, not only against mocks.
- In week 9, ready consumers move to v2 while the provider keeps 1 canonical v2 business-logic path and translates v1 at the boundary.
- v1 and v2 remain supported together for at least 18 weeks after the v2 launch: a full 16-week migration window followed by 14 consecutive days below 1 percent v1 traffic and no approved exception, so the earliest v1 shutdown is the end of week 27.
Why interviewers ask this: The interviewer is checking whether you can trade a quantified compatibility cost for schedule certainty without creating an unbounded legacy API.
I would use expand-and-contract because 30 days of mixed application versions makes a big-bang customer schema change unsafe.
- In week 1, data owners freeze invariants, source-of-truth rules, nullable behavior, and reconciliation thresholds in a versioned customer schema contract.
- Week 2 adds backward-compatible columns and readers; dual writes begin only after shadow comparisons stay below 0.01 percent mismatch for 24 hours.
- A throttled backfill runs at no more than 20,000 rows per second, and all 6 teams move reads by week 8 against the same compatibility matrix.
- Old fields are removed no earlier than week 12 and only after the 30-day coexistence period ends, 7 consecutive days show zero old reads, and rollback and reconciliation gates pass.
Why interviewers ask this: The interviewer is evaluating whether you can sequence a large customer schema migration around compatibility, measurable correctness, and a finite rollback window.
I would carve out the 3-feature launch slice while preserving a direct path into the shared feature store.
- By day 3, the source and consumer teams freeze the 3 feature definitions, a 5-minute freshness limit, ownership, retention, and acceptance tests.
- A producer-owned append-only feed lands in week 3 using the same feature identifiers and types that the final store will ingest, so it does not create a second computation model.
- Checkout integrates during weeks 4 to 6; the week-6 gate requires 99.9 percent feed availability and less than 0.1 percent value mismatch over 7 days.
- The full store shadows the feed for 14 days after its week-10 delivery, and checkout switches by week 12 only if parity passes; otherwise the bounded feed expires in week 14 and forces a scope or date decision.
Why interviewers ask this: The interviewer is checking whether you can remove an oversized technical dependency from the critical path without creating a permanent parallel platform.
I would break the cycle with minimum contracts and a temporary adapter rather than schedule a fragile simultaneous handoff.
- A 2-day sequencing workshop identifies the smallest stable surfaces: consent event v1, the profile read schema, and the auth SDK interface.
- By day 3, auth builds against the frozen profile schema while mobile emits consent event v1 independently of the final profile implementation.
- A platform-owned adapter converts the current consent payload until the end of week 6; native producers must be ready by week 5, followed by 7 days with zero adapter calls before removal.
- If the week-2 end-to-end gate fails, I remove offline profile enrichment and recover 10 working days instead of preserving the cycle through a coordinated release.
Why interviewers ask this: The interviewer is evaluating whether you can detect a dependency cycle and remove it through explicit contracts, bounded compatibility work, and a quantified scope choice.
I would combine contract-first mocks with a thin production slice because mocks alone create false schedule progress.
- Within 24 hours, the provider and consumers classify 3 endpoints as launch-critical and move the other 8 beyond the first release boundary.
- An OpenAPI contract and generated mock land by day 2, but consumer readiness remains provisional until the same tests pass against the provider build.
- The provider delivers token validation, user lookup, and revocation by week 5, while consumer teams resequence UI and local validation work around that gate.
- If the week-5 integration gate misses, I remove account linking to recover 2 weeks or move the date instead of adding overtime across 7 teams.
Why interviewers ask this: The interviewer is checking whether you distinguish useful contract-first parallelism from mock-based optimism and preserve a dated fallback decision.
I would remove the global-ordering dependency after proving that the domain requires ordering only within each account.
- Within 2 days, domain owners enumerate ordering invariants; any required cross-account invariant keeps the 12-week broker path and moves the launch by 3 weeks.
- By day 4, the event contract includes account ID, a monotonic account sequence, event ID, and a 24-hour replay window.
- The broker delivers the partitioned stream in week 3, and all 6 teams integrate by week 6 against tests at 2 million events per minute and 4 million-event bursts.
- A week-7 replay of 100 million events must show zero per-account inversions and fewer than 0.01 percent duplicates, all reconciled, leaving 2 weeks for launch hardening.
Why interviewers ask this: The interviewer is evaluating whether you can challenge an unnecessary ordering dependency and replace it with explicit technical invariants and measurable gates.
I would model the lab as a constrained dependency and reserve slots by schedule impact, not request order.
- The integrated graph includes lab duration, setup time, retest probability, earliest readiness, and latest safe slot for all 3 paths.
- Each week, I reserve 12 hours for the critical path, 5 hours for the near-critical path, and 3 hours for failed-test recovery, then recalculate from actual results.
- Teams enter the lab only after emulator, contract, and fixture checks pass at least 95 percent, keeping basic defects out of the shared bottleneck.
- Full-chain gates in weeks 5 and 8 decide between cutting the lowest-volume device matrix, which saves 4 days, and changing the certification date.
Why interviewers ask this: The interviewer is checking whether you can incorporate a scarce technical resource into the dependency model and make a quantified recovery choice.
I would version the event and fund a bounded translator because an atomic switch cannot fit both consumer schedules and an in-place semantic change is ambiguous.
- Within 48 hours, the teams define currency, precision, rounding, idempotency, and replay behavior in a v2 schema with compatibility tests.
- The fast consumer passes its gate in week 3; at the week-5 launch, the producer emits canonical v2 and the translator supplies v1 to the slow consumer.
- Before week 4, a replay of 10 million events must show zero price divergence, sustain 1 million events per minute, and add no more than 15 percent to p99 latency.
- The slow consumer passes its gate in week 7, and the translator remains funded until 14 consecutive days below 1 percent v1 traffic, making week 9 the earliest retirement date.
Why interviewers ask this: The interviewer is evaluating whether you can manage a high-throughput semantic contract change through versioning, replay evidence, and a bridge that covers the full 7-week migration.
I would use 2 waves because parallel migration of all 8 teams would expose the shared dependency and consume the 4-day float at once.
- The adapter and authorization contract complete in weeks 1 to 5, with rollback behavior and policy equivalence proven before any consumer changes production traffic.
- Two teams with different access patterns migrate in weeks 6 and 7; wave 2 starts in week 8 only after 5 consecutive days with zero contract violations across at least 100,000 authorization decisions.
- The remaining 6 teams migrate during weeks 8 to 13, followed by final validation in weeks 14 and 15 and the production handoff on day 1 of week 16.
- A gate slip of up to 4 working days consumes the stated float; a larger slip defers the 3 lowest-volume teams instead of compressing the 2-week validation period.
Why interviewers ask this: The interviewer is checking whether you can build mathematically consistent waves around a near-zero-float critical path and preserve an explicit fallback.
I would sequence the roadmap around the smallest platform capabilities that unlock later outcomes, not divide the shortfall evenly across teams.
- I map each requested capability to its Q2 to Q4 consumers, required-by date, acceptance evidence, and scarce skills, then identify the dependency chains that control value.
- I fund Q1 contract stabilization and migration tooling if they unlock four product teams, while deferring isolated dashboard work with no downstream dependency.
- I reserve 15 percent of the 420 engineer-weeks for integration, technical risk, and forecast variance instead of planning all capacity to feature scope.
- The published roadmap separates Q1 commitments from Q2 forecasts and names the evidence that will promote each later item into a commitment.
Why interviewers ask this: The interviewer is evaluating whether the candidate can turn constrained capacity and downstream dependencies into a credible multi-quarter sequence.
I would mark the milestone red because consumer readiness and production evidence, not implementation effort, determine migration progress.
- I replace percent complete with three exit measures: 12 signed contracts, 12 passing contract suites, and staged production traffic with error and latency thresholds.
- The provider and each consumer receive dated gaps, owners, and a required migration wave, starting with two representative consumers within 14 days.
- I stop nonessential provider features until the first wave proves compatibility, telemetry, and rollback under real traffic.
- If fewer than 8 consumers pass by the next monthly checkpoint, I replan the shutdown date and funded dual-support period rather than preserve the quarter label.
Why interviewers ask this: The interviewer is checking whether milestones are governed by verifiable integration evidence and explicit replanning triggers.
I would sequence the identity work by hard deadline, shared enablement, and cost of delay, then expose what the team cannot do concurrently.
- I first determine whether token modernization is a prerequisite for residency or SSO; if it is, I fund the thinnest production-ready slice that unlocks both.
- EU residency receives the protected path to September 30 because the regulatory date is fixed, including review lead time and regional validation rather than code completion alone.
- SSO follows in vertical increments for the highest-value customers, with the $4 million pipeline discounted by close probability and customer required-by dates.
- I publish one capacity-backed sequence with displaced scope, decision dates, and a contingency if modernization misses its Q2 proof milestone.
Why interviewers ask this: The interviewer is evaluating sequencing across a fixed obligation, commercial value, and an enabling technical investment.
I would hold the date only with a smaller safe scope that meets the 700 millisecond target and has a tested rollback path.
- I identify the minimum end-to-end checkout path, moving optional recommendations and two low-volume payment methods behind disabled flags.
- The launch gate requires representative peak load at 1.5 times forecast traffic, p99 below 700 milliseconds, error rate below 0.5 percent, and a completed rollback drill.
- I compare the reduced launch with a date move using revenue exposure, customer coverage, support load, and the cost of restoring deferred scope.
- If the performance gate fails by the last reversible date, I recommend delay rather than silently spend reliability to protect the calendar.
Why interviewers ask this: The interviewer is checking whether the candidate can make a quantified scope, date, and reliability decision under a real market deadline.
I would report a forecast range and confidence level grounded in current evidence rather than average the five team dates into one promise.
- I model dependency-level ranges, correlate the two shared risks, and use the 20 to 35 percent historical error to calibrate rather than pad every task independently.
- I present a P50 date for internal sequencing and a P80 date for an external commitment, with the assumptions that separate them.
- Each unresolved estimate gets an evidence event, such as a scale test or design approval, plus the date when it should narrow the forecast.
- Confidence changes only when throughput, dependencies, scope, or technical evidence changes, and every update explains the movement from the previous range.
Why interviewers ask this: The interviewer is evaluating quantitative roadmap confidence, calibration, and the distinction between forecast and commitment.
I would narrow the release train to work that truly needs synchronized certification and let independently deployable services release continuously.
- I segment the missed items by shared integration, approval wait, environment contention, and simple scope lateness to test whether the train solves the actual constraint.
- For coupled changes, I keep a monthly integration window with entry evidence seven days earlier and allow incomplete scope to leave without blocking the train.
- For decoupled services, automated contract tests, progressive delivery, and local rollback authority replace the common date.
- After two quarters, I keep the model only if miss rate falls below 15 percent and median completed-work wait falls below five days without higher change failure.
Why interviewers ask this: The interviewer is checking whether the candidate can redesign a release train from measured coordination value and batching cost.
I would treat the 40 engineer-weeks as a measured delivery tax and fund debt items that demonstrably return capacity or reduce roadmap risk.
- I rank debt by recurring toil, rework, reliability exposure, blocked features, and time to payback instead of accepting one undifferentiated backlog.
- I fund automation expected to remove 12 toil weeks per quarter and a dependency upgrade required by two committed launches, with explicit acceptance measures.
- Product scope is reduced to fit remaining capacity, while small debt fixes stay embedded in related feature work where that avoids a second migration.
- The quarterly review compares realized capacity returned and avoided rework with the promised baseline, stopping debt work that produces no measurable change.
Why interviewers ask this: The interviewer is evaluating whether technical debt competes through quantified carrying cost, roadmap enablement, and verified payoff.
I would pause broad investment and run one bounded replan only if evidence shows the adoption barrier can be removed within a quarter.
- I separate product gaps, migration cost, reliability, and weak demand by interviewing the 3 adopters and a representative sample of the 17 non-adopters.
- The replan gets 8 engineer-months to reduce setup below eight hours and bring five additional teams onto real production workloads within six weeks.
- I preserve reusable automation but stop feature expansion and new platform dependencies during the proof period.
- If either threshold fails, I kill the platform, support the three adopters through an exit plan, and return remaining capacity to higher-value roadmap work.
Why interviewers ask this: The interviewer is checking whether the candidate can overcome sunk-cost bias with specific, time-bound continue and kill criteria.
I would rebaseline the three milestones around a tested alternative, not leave the vendor date as an unmanaged external dependency.
- In the first two days, I confirm the contractual commitment, technical gap, downstream required-by dates, and which milestone outcomes survive without the capability.
- Engineering time-boxes a compatibility layer and a second-vendor proof, measuring performance, migration effort, data portability, and cleanup cost.
- By day seven, I present delay, reduced scope, workaround, and vendor-switch options with confidence, spend, reliability, and reversibility.
- By day ten, the selected path has funded owners, revised milestone evidence, a vendor exit trigger, and updated Q3 to Q4 commitments for every consumer.
Why interviewers ask this: The interviewer is evaluating rapid roadmap replanning around a material vendor dependency with technical and commercial evidence.
I would remove complete outcomes and concurrency rather than reduce every workstream by 20 percent and make all milestones unreliable.
- I protect regulatory commitments, reliability floors, and enabling work already on the critical path, then rank remaining outcomes by cost of delay and dependency leverage.
- I stop one low-adoption self-service workstream, defer a regional expansion, and keep the migration slice that retires the largest legacy operating cost.
- The new baseline shows reclaimed engineer-weeks, changed consumer dates, residual technical debt, and one accountable decision owner for each displaced outcome.
- I reopen the plan only if capacity returns, adoption exceeds the agreed threshold, or milestone evidence changes the value or risk assumptions, not because a sponsor escalates again.
Why interviewers ask this: The interviewer is checking whether the candidate can make a coherent roadmap cut and define evidence-based criteria for later replanning.
Locked questions
- 21
Sixty services need unified observability within six months: a vendor quotes $900,000 per year, while engineering estimates six engineers for two quarters to build internally. How would you facilitate the build-versus-buy decision?
procurementobservabilityestimation - 22
A SaaS platform hosts 200 enterprise tenants in one deployment costing $180,000 per month. A dedicated stack adds $22,000 per tenant per month, but three regulated prospects require demonstrable isolation. How would you facilitate the shared-versus-dedicated tenancy decision?
deploymentcloud - 23
A global inventory service handles 20,000 reservations per second across three regions. Product allows at most 0.01% sales beyond available stock, while engineering says strong consistency would raise p99 latency from 120 ms to 280 ms. How would you facilitate the consistency decision?
consistencylatency - 24
Checkout needs a fraud decision within a 300 ms budget, but the fraud service reaches 2 seconds at peak and is available 99.7%. How would you guide the synchronous-versus-asynchronous integration choice?
async - 25
A public API has 80 known consumers, 15% of traffic has no identified owner, and a breaking schema change must ship within nine months. How would you lead the versioning decision?
versioningschema - 26
7 product teams are building separate entitlement services. A shared platform needs 6 engineers for 2 quarters, while each team can ship its local service in 8 weeks. How would you facilitate shared-platform versus team-owned ownership?
ownership - 27
Analytics processes 2 TB per day in a four-hour batch for $18,000 per month, while stakeholders request five-minute freshness and streaming is estimated at $55,000 per month. How would you lead the batch-versus-streaming decision?
stakeholder-managementestimationcommunication - 28
A service for eight million users is expanding to three regions. Active-active adds $140,000 per month and 35 ms to write p99, while active-passive has a tested 20-minute recovery objective. How would you shape the rollout architecture decision?
releasesarchitecture - 29
A search API serves 12,000 requests per second at 180 ms p99 and 99.95% availability. Reaching 120 ms and 99.99% is estimated to cost another $900,000 annually. How would you broker the cost, latency, and reliability trade-off?
estimationapilatency - 30
A legacy order platform handles three million requests per day and costs $90,000 per month. A big-bang replacement needs a six-hour write freeze, while a strangler rollout adds $60,000 per month for six months. How would you lead the rollout choice?
releasesmigration - 31
Seven teams are adding a cross-service order workflow across five services at 3,000 orders per second, and retries could duplicate charges or lose inventory updates; how do you reduce correctness risk before the eight-week launch?
risk-managementlaunches - 32
You must migrate 18 TB and 2.4 billion records during a 90-minute cutover window across five data-owning teams; what risk plan do you require before approving the date?
risk-management - 33
A checkout request crosses six services and must stay below 350 ms p95 and 800 ms p99 at 4,000 RPS; how do you establish and govern latency budgets between teams?
latency - 34
A new payments flow has a 99.95% request-success SLO over 28 days and forecasts 200 million requests; how do you use its error budget to control rollout risk before launch?
risk-managementlaunchesreleases - 35
Traffic is forecast to grow from 6,000 to 24,000 RPS in one quarter, and leadership requires 35% capacity headroom with p99 below 500 ms; how do you build the capacity risk plan across five teams?
risk-managementcapacity-planningcapacity - 36
Eight teams are preparing a three-region rollout with a 99.99% availability target, 15-minute RTO, and 60-second RPO; how do you sequence the program before the first customer wave?
program-managementreleases - 37
An 11-team payment program launches in 16 weeks across six jurisdictions and depends on PCI-DSS evidence, GDPR transfer reviews, and three security architecture changes; how do you keep compliance from becoming a late blocker?
program-managementlaunchesarchitecture - 38
A rollout changes schemas and APIs across 14 services while 1.2 billion records are backfilled over 21 days; what rollback and readiness evidence do you require before go-live?
health-checksrollbackschema - 39
A critical identity vendor must support 12 services, three regions, and 8,000 logins per second in ten weeks; what prelaunch risks and decisions do you drive?
procurement - 40
Ten teams and 25 services are four weeks from a launch expected to process two million transactions per day; what readiness package is sufficient for a go or no-go decision?
launchestransactionshealth-checks - 41
A VP asks you to rank four teams next Friday using their last six sprint velocities of 42, 31, 68, and 27 points; do you publish the ranking, normalize the points, or replace the comparison?
agilenormalization - 42
A three-team developer onboarding flow has a 24-day lead time but only 8 days of active cycle time, and the target is 14 days within 10 weeks; do you add 2 engineers, automate intake, or cut work in progress?
onboarding - 43
Six weeks before a cutover, directors call the program 80 percent ready, but only 21 of 40 readiness criteria have current evidence and 3 of 8 critical gates lack evidence; do you report 80 percent, 53 percent, or no percentage?
program-managementhealth-checks - 44
The CTO wants a DORA scorecard for 18 teams in 8 weeks and proposes a league table with a daily-deployment target; do you ship that table or a contextual team-owned scorecard?
deployment - 45
A 12-program portfolio reports 92 percent milestone completion, yet only 7 of 12 programs have accepted outcomes and 4 of the other 5 are more than 30 days late; which metric set do you put in the monthly review in 2 weeks?
milestonesprogram-managementplanning - 46
Over the next 6 calendar weeks, SRE has 12 engineer-weeks available: launch A requests 7 engineer-weeks, has a fixed date in 5 calendar weeks and $2 million contractual exposure; launch B requests 10, targets 7 calendar weeks, carries $500,000 expected value, may slip 2 calendar weeks, and needs 5 engineer-weeks before an external review at the end of week 6; launch C requests 8, targets 9 calendar weeks, has $100,000 annual savings, no external dependency, and a flexible date. All estimates are p80 and none of the work has a qualified substitute. How do you prioritize and allocate the 12 engineer-weeks?
schedulingdependencieslaunches - 47
Six backend engineers provide 60 engineer-weeks across 10 calendar weeks: a contractual platform migration needs 32 engineer-weeks, a revenue feature needs 26, and 10 engineer-weeks must remain for support; do you split the engineers 3 and 3, pause the feature, or cut its scope?
scope-managementmigrations - 48
Four data engineers must process 24 production model approvals in 4 weeks, but only 2 are certified reviewers and each can approve 2 changes per week after operations work. One other engineer can qualify after 4 paired approvals completed by the end of week 2, and then can approve 2 per week. Demand includes 6 compliance changes due in week 2, 10 customer commitments due in week 4, and 8 flexible internal changes. What capacity plan do you commit?
capacity-planningcapacityconcurrency - 49
Eight engineers provide 104 gross engineer-weeks in a 13-week quarter, the last 4 quarters show unplanned work of 18, 22, 20, and 20 percent, fixed leave and governance overhead is another non-overlapping 15 percent, and the roadmap requests 82 engineer-weeks. What available-capacity forecast and commitment ceiling do you publish?
capacity-planninggovernancecapacity - 50
Fourteen teams are deadlocked on a cross-org RFC for 20,000 billing updates per second with a 5-minute freshness SLO, and a decision is due in 15 business days; do you choose REST fan-out, Kafka events, or continue both designs?
kafkadesignfan-out - 51
Eight teams are 7 weeks from a warehouse-control rollout to 12 sites, but the integrated plan now needs 11 weeks and a full delay puts $18 million of seasonal throughput at risk. Do you defend the original scope, move the date, or rescue a smaller launch?
risk-managementscope-managementlaunches - 52
You inherit a 10-team claims-engine program that is 14 months old, has missed 3 milestones, spent $9.6 million of a $12 million budget, and matches the legacy result on only 82% of claims against a 99.95% target. A regulatory rule takes effect in 16 weeks. Do you reset, continue, or close the program?
milestonesprogram-managementplanning - 53
Seven teams have spent 2 quarters and 84 engineer-months rebuilding a video-transcoding platform. A public launch is 10 weeks away, but cost is $0.027 per minute against a $0.012 goal, p95 completion is 11 minutes against 4, and the incumbent costs $0.018. Do you kill or continue the rebuild?
launches - 54
Six teams are 4 weeks from a board deadline to remove static production credentials, but only 41 of 130 services have migrated, the program is already 5 weeks late, and security found 17 exposed keys in the last 30 days. Do you stop, continue unchanged, or narrow the work?
program-managementestimation - 55
Nine teams have 8 weeks to release enterprise device management for Windows, macOS, Linux, 6 identity integrations, and 2 customers representing $6 million in annual contracts. Integrated testing forecasts 13 weeks, while Windows, macOS, and 2 integrations can be ready in 7. Do you rescope or move the date?
releasestesting - 56
Seven teams are 18 days from sending a new content-moderation model to 20% of traffic. Product wants the launch, engineering wants a 5% canary, and security claims veto authority because false negatives are 2.4% against a 1.0% limit. How do you resolve the authority conflict under the deadline?
launchesestimationdeployment-strategies - 57
Six teams need a shared shipment-ETA API for a retailer launch in 5 weeks, but logistics and commerce each refuse ownership. The contract changed 14 times last quarter, p99 latency is 900 ms against 300 ms, and 7 incidents in 60 days had no clear responder. What ownership decision do you drive?
incidentsownershipapi - 58
Twenty-four teams have 14 weeks to open an automated baggage system, with 63 cross-team dependencies, 19 items marked critical, and 17 handoffs lacking acceptance tests. The airport date cannot move without a $4 million penalty. How do you turn the network into an executable priority decision?
acceptancesystem-designdependencies - 59
Twenty-two teams must leave a container runtime that loses security patches in 9 weeks. Only 58 of 146 services have moved, the remaining 88 carry 72% of production traffic, and platform on-call can support both runtimes for only 6 more weeks. Do you grant a broad extension or force the migration?
on-callmigrationscontainers - 60
Five teams have 12 weeks to migrate a 9 TB product catalog with 18,000 writes per second and no more than 0.01% record divergence. Eleven application write paths cannot all be changed before week 10, while dual writes add 35 ms to a 120 ms p99 budget. Do you choose application dual writes or log-based change-data capture?
- 61
Eight teams are 10 weeks from a three-region launch with a 99.99% availability SLO. Legal says EU customer data cannot leave the EU, while the current active-active design replicates it to the US and costs $240,000 per month. The CRO insists all regions launch together. Do you keep the date, change the topology, or remove the EU?
designsloreplication - 62
A payment rollout has reached 25% of traffic across six teams with no Sev-1, but error-budget burn is 6x, p99 latency is up 42%, and ledger drift is 0.08% against a 0.01% limit. The launch owner wants to advance to 50% tonight to hold a date in three days. Do you pause, roll back, or continue?
launchesreleasesrollback - 63
A change gate used by 22 teams has increased median lead time from three to nine days, 35% of changes bypass it, and only one of the last 40 failed changes was caught by the manual review. A regulated launch is 12 weeks away. Do you abolish the gate, enforce it harder, or rebuild it while delivery continues?
launches - 64
Sixteen engineering teams refuse a DORA scorecard after a VP used an early dashboard to rank them, data coverage is only 72%, and the CTO still wants a 30% lead-time improvement this quarter. You have eight weeks to recover adoption. Do you mandate reporting, abandon the metrics, or reset the program?
program-managementcoveragedecision-making - 65
A Sev-1 outage spans 14 services and nine teams, but six service-catalog ownership records are stale and responders have already lost 12 minutes paging the wrong teams. The affected platform has a 99.95% SLO and a 30-minute mitigation target. Do you repair ownership first or change incident command immediately?
incidentsownershipslo - 66
A shared checkout outage involves 11 teams from three organizations that use Sev-1, P0, and Critical labels with different command rules. Revenue loss is $120,000 per minute, p99 has risen from 400 ms to 4.5 seconds, and no one has authority after eight minutes. Do you reconcile the frameworks during the outage or impose one command model?
- 67
At 8:00 a.m. you confirm that a nine-team flagship launch in five days will slip at least six weeks: peak tests show 99.2% success against a 99.95% SLO, and a $12 million campaign starts in 48 hours. The CEO expects an update in 90 minutes. Do you protect the date with reduced scope or tell the CEO to move it?
scope-managementlaunchestesting - 68
Four days before a seven-team enterprise launch worth $8 million annually, the CISO blocks release because privileged admin logs are editable and retained for only 30 days instead of the required immutable 365 days. Engineering calls exploitation theoretical and estimates a three-week fix. Do you seek an exception, cut scope, or delay?
scope-managementlaunchesreleases - 69
Discovery raises a five-team platform program from $12 million to $15.6 million, a 30% funding gap, after $6 million has already been spent. The expected annual savings remain $8 million, but a scale assumption is still unproven, and the CFO decides in 10 business days. Do you request all $3.6 million, rescope, or stop?
program-management - 70
Two platform outages of 47 and 83 minutes in the last 30 days reduced availability to about 99.70% against a 99.95% SLO and triggered $4.8 million in SLA credits. Eleven teams are still shipping roadmap work, and the board meets in 72 hours. Do you ask for a broad feature freeze, a funded reliability plan, or both?
roadmapsloroadmapping - 71
Three engineering directors are split 6 weeks before platform budget lock: renew the managed PaaS for $1.7 million, spend $3.8 million migrating 26 services to Kubernetes, or fund both for $5.5 million against a $4.2 million cap. One regulated launch cannot run on the PaaS. What strategy do you recommend when none of the directors reports to you?
kubernetescloudlaunches - 72
An approved API-versioning RFC gives 24 teams 16 weeks to stop shipping unversioned endpoints, but 7 teams insist they need 24 weeks and a $3 million partner launch depends on 2 of them. The platform team can maintain only 1 temporary compatibility layer. How do you decide the adoption plan?
endpointsversioningdecision-making - 73
A cross-org RFC is rejected after 10 weeks of work because 5 of 17 affected teams were consulted only at final review, and the architecture freeze is 21 days away. Do you defend the current proposal, withdraw it, or restart the decision from zero?
decision-makingarchitecture - 74
Two programs need the same platform team over the next 12 weeks, but only 32 engineer-weeks are available. Regulatory remediation needs 20 engineer-weeks before a fixed deadline with $8 million exposure; a growth launch needs 24, but a smaller cohort preserving $2 million of its $6 million ARR needs 12. What portfolio decision do you drive?
program-managementlaunchesestimation - 75
Annual funding is fixed at $9 million: regulatory remediation due September 30 requires $4 million, reliability work after 3 consecutive months at 3 times the error-budget burn needs $3 million, and a growth launch projected to add $12 million ARR requests $5 million. How do you allocate the money and decide what ships?
launchesprioritization - 76
The observability stack for 55 services reaches end of support in 14 weeks. An internal replacement needs 24 weeks and $2.4 million in its first year; a vendor can be live in 10 weeks for $1.6 million annually but wants a 3-year commitment. Do you build, buy, or pay $900,000 for a 6-month extension of the old stack?
build-buyobservabilityprocurement - 77
Five weeks before a $4 million customer launch, procurement for the identity vendor is 3 weeks late: privacy needs a data-transfer addendum, legal rejects the liability cap, and finance has not approved $650,000 of annual spend. A reduced launch without enterprise SSO remains possible. What decision sequence do you set?
procurementlaunches - 78
Eight weeks into a 14-week rollout, a strategic vendor is acquired: availability falls from 99.98% to 99.4%, support response rises from 2 to 19 hours, 40% of workloads have migrated, and your termination window closes in 30 days. A second provider can take traffic in 6 weeks for $350,000. What do you decide?
procurementreleases - 79
A mid-level TPM asks you to take over a 9-team developer-platform program with 12 weeks remaining, 60 engineer-weeks already allocated, and 3 sponsors separately demanding cost reduction, faster onboarding, and policy compliance. No shared success measure exists. How do you mentor them and decide whether implementation continues?
mentoringonboardingsponsor - 80
A mid-level TPM has produced 6 green weekly reports but missed 3 technical risks that caused a 2-week slip; the next phase launches in 7 weeks against a $5 million customer commitment and spans 12 services. Do you keep them on the critical path, reduce their scope, or replace them?
schedulingdependenciesscope-management - 81
You must hire 12 senior TPMs in 8 weeks, but across the last 25 loops in each region, pass rates are 52% in the Americas, 24% in Europe, and 60% in Asia. How do you correct the inconsistency without missing the deadline or lowering the bar?
estimation - 82
A TPM wants promotion to senior in the cycle 10 weeks away. They delivered 4 launches on time, but partner feedback says their manager resolved the hardest cross-team decisions and they have not created a reusable operating mechanism. What promotion plan do you set?
launchesfeedbackcross-team - 83
A program spanning 25 teams now has 14 recurring forums consuming 220 person-hours per week, while cross-team decisions still wait a median of 9 days. How do you redesign the operating model in the next month?
program-managementcross-team - 84
Eight weeks before a global launch, 43% of handoffs between Americas, Europe, and Asia are reopened and blocked work waits a median of 19 hours. What execution change do you make?
launches - 85
Half of the teams in a 40-service platform migration move to new organizations with 12 weeks remaining, and both executive sponsors are replaced. How do you rebaseline without restarting the program?
program-managementmigrationssponsor - 86
After an acquisition, 90 services must support one customer experience in 6 months, but 37 have no confirmed owner and the companies use incompatible release approvals; 14 changes for the first joint release have already waited a median of 6 days, and that release is 8 weeks away. Where do you start?
releases - 87
A Sev-1 starts 36 hours before a critical product launch, checkout errors reach 12%, and 5 of the 8 launch-readiness engineers are the only qualified responders. How do you lead both efforts?
launcheshealth-checks - 88
All 18 milestones on a rollout dashboard are green, but customer adoption fell from 38% to 21%, priority incidents rose from 3 to 7 per month, and cost per transaction increased 26%. What status and next decision do you set?
milestonesplanningreleases - 89
The CEO publicly promises a full global launch in 9 weeks before engineering estimates it; the first technical review produces a P50 of 13 weeks, a P80 of 16, and 4 unresolved critical dependencies. What do you do in the next 72 hours?
launchespromisesestimation - 90
Two principal architects are deadlocked over a centralized API gateway versus service-mesh sidecars for enforcing policy on 14 teams and 80,000 requests per second; teams have stopped for 8 days, the added p99 budget is 12 milliseconds, and the compliance launch is 10 weeks away. How do you break the deadlock?
launchesapilocking - 91
Three weeks before an 8-service identity cutover, tracing reveals that 38 percent of authorization calls still use an undocumented legacy profile lookup. The replacement fails 6 percent of requests at the forecast peak of 18,000 RPS, against a 99.95 percent SLO, and extending the legacy contract costs $180,000. Do you repair, bridge, or delay?
authslo - 92
Ten days after a reorganization transfers a 12-service token platform, the receiving team has missed 2 pages for 22 minutes and cannot perform a rollback; a $7 million migration depends on the platform in 6 weeks and its availability SLO is 99.99 percent. Do you keep the organizational handoff date or restore shared ownership?
ownershipmigrationstokens - 93
Five weeks before an EU analytics launch, the privacy review blocks 2.4 billion raw events per day from being copied to a US lake because 7 of 24 fields can identify users. The launch supports $6 million in annual contracts, while GDPR exposure could reach €20 million; how do you redirect the program?
program-managementlaunchesgdpr - 94
Six weeks before a payment launch tied to $8 million in contracts, the PCI assessor finds card numbers touching 11 services and says the current control plan needs 9 more weeks. Authorization must meet a 99.95 percent availability SLO and 800 millisecond p99 latency; what schedule and scope do you commit to?
scope-managementlaunchesauth - 95
Forty-two minutes after a database cutover at 30,000 writes per second, reconciliation shows 0.7 percent of writes split between the old and new systems, roughly 529,000 records. The service carries $2.5 million in daily transactions and a 99.99 percent availability SLO; do you roll back, repair forward, or keep ramping?
databasetransactionssystem-design - 96
Ten days before 3 launches, two Sev-1 incidents remove 5 of 8 SREs for the next 14 days. Payments has burned 60 percent of its monthly error budget, a deletion service has a regulatory deadline in 12 days, and a search launch supports a $3 million campaign; how do you allocate the 30 remaining SRE engineer-days?
incidentsestimationreliability - 97
Six weeks and $900,000 into a recommendation program, finance rejects the $1.8 million event platform it requires; the remaining recommendation work costs $4.1 million, and without the platform coverage falls from 100 to 20 percent of traffic while forecast annual value drops from $9 million to $1 million. What do you fund now?
program-managementcoverage - 98
A VP overrides your recommendation to stop at 25 percent after a canary burns error budget at 3.2 times the allowed rate. The full rollout then causes a 47-minute outage, 180,000 failed checkouts, and $2.1 million in credits; what decisions do you make during and after the incident?
reliabilitydeployment-strategiesincidents - 99
In an 18-team, $35 million portfolio, 14 technical decisions are waiting for you, median decision age has reached 9 days, and two 99.95 percent services have missed release gates by 3 weeks. How do you remove yourself as the bottleneck without losing control?
trackingreleasesportfolio - 100
A 6-week rescue saves a $12 million contract, but 74 people averaged 58-hour weeks, 31 emergency decisions bypassed normal review, and the launch produced 2 Sev-2 incidents against a 99.95 percent SLO. The CEO wants phase 2 to start in 5 days; what do you do?
launchesincidentsslo