FinOps Engineer interview questions
100 real questions with model answers and explanations for Senior candidates.
See a FinOps Engineer resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would establish a federated FinOps practice with a small central team and accountable owners inside engineering, finance, and product.
- The central team would own cost data, commitment strategy, policy, and enablement, while 40 team delegates own workload decisions.
- Finance would own budgets and accruals, engineering would own usage efficiency, and product would own unit-economics targets.
- I would sequence visibility in months 1 to 3, optimization in months 4 to 6, and governed automation in months 7 to 12.
- Success would mean at least 95% allocated spend, forecasts within 5%, and verified savings rather than dashboard activity.
Why interviewers ask this: The interviewer is testing whether the candidate can turn the FinOps Framework into clear ownership, phases, and measurable outcomes at enterprise scale.
I would define persona-specific decisions and give each persona a self-service view rather than route every question through five specialists.
- Engineers get workload cost, efficiency recommendations, and owner-level alerts tied to deploy and service metadata.
- Engineering managers get budget variance, commitment coverage, and unit cost for their portfolio of services.
- Finance gets amortized cost, forecast, accrual, and variance views that reconcile to provider invoices.
- Executives get trend, risk, and business-unit economics, while a network of trained champions handles first-line questions across 600 engineers.
Why interviewers ask this: The interviewer is evaluating whether personas translate into distinct decisions, data products, and scalable support boundaries.
I would assign one accountable owner per financial decision and keep the FinOps team responsible for analysis and controls, not every approval.
- Cost-center leaders are accountable for budgets and variance, with finance responsible for the planning calendar and ledger mapping.
- Application owners are responsible for allocation metadata, while FinOps owns allocation rules and reports coverage below 98%.
- FinOps recommends commitments, procurement owns contract execution, and the CFO approves exposure above the agreed $5M threshold.
- Engineering leaders approve workload changes against SLA, while FinOps verifies realized dollars after implementation.
Why interviewers ask this: A strong answer separates accountability from execution and prevents FinOps from becoming the owner of decisions that belong to finance or engineering.
I would mature the capabilities in dependency order and limit each 6-month wave to outcomes the organization can operate.
- Wave 1 establishes billing ingestion, account ownership, 90% allocation, budgets, and a monthly forecast.
- Wave 2 adds anomaly governance, rightsizing workflow, commitment management, and showback for the largest 10 teams.
- Wave 3 introduces chargeback, unit economics, policy automation, and savings verification after the data controls are stable.
- Each capability advances only when an owner, cadence, control metric, and 3 months of reliable operation are demonstrated.
Why interviewers ask this: The interviewer is checking whether maturity is treated as an operational sequence with exit criteria rather than a certification checklist.
I would push routine decisions into asynchronous workflows and reserve leadership time for material exceptions.
- Daily pipelines refresh allocation, unit cost, commitment utilization, and anomalies without requiring meetings.
- Teams receive a weekly digest with only actions above $10K annualized value or 10% budget variance.
- A 45-minute monthly review covers forecast, realized savings, and blocked actions for each business unit.
- A quarterly 60-minute executive review handles commitment exposure, product economics, and policy decisions across the $50M portfolio.
Why interviewers ask this: The interviewer is evaluating whether governance cadence is proportional to decision value and respects limited leadership attention.
I would balance financial outcomes, data trust, commitment risk, and engineering adoption instead of using gross savings alone.
- Financial KPIs are realized savings, unit-cost movement, forecast variance, and budget variance, all measured against a fixed baseline.
- Data KPIs are allocation coverage above 98%, invoice reconciliation within 0.5%, and feed freshness under 24 hours.
- Commitment KPIs separate coverage, utilization, break-even risk, and expirations due within 120 days by cloud.
- Adoption KPIs track action closure, policy exceptions, and the share of 12% improvement delivered without breaching SLA.
Why interviewers ask this: The interviewer is testing whether the scorecard distinguishes real economic outcomes from activity metrics and discounted-spend vanity metrics.
I would preserve each provider feed raw, normalize it into a FOCUS-aligned canonical layer, and publish governed marts for finance and engineering.
- Immutable raw tables retain invoice IDs, corrections, pricing fields, and source lineage so transformations remain auditable.
- dbt models map AWS CUR, Azure exports, and GCP detailed billing into shared billing account, resource, service, charge, and commitment dimensions.
- Snowflake clustering and incremental loads keep the 2 billion-row monthly pipeline inside its 6-hour freshness target.
- Semantic marts expose amortized cost, allocation, commitments, and unit economics without forcing users to interpret provider-specific columns.
Why interviewers ask this: The interviewer is evaluating layered data architecture, FOCUS normalization, lineage, and performance at billing-scale volume.
I would centralize CUR 2.0 exports in a dedicated billing data account and process them incrementally from S3.
- Management-account exports include resource IDs, split cost allocation data, Savings Plan, RI, credit, and amortization fields.
- S3 event manifests trigger idempotent loads keyed by billing period, account, and file checksum rather than rescanning all history.
- A raw-to-staging job records schema drift and quarantines malformed files while allowing unaffected accounts to meet the 8-hour target.
- Reprocessing a closed month creates a new version so late credits and refunds remain traceable to the $35M invoice base.
Why interviewers ask this: A strong answer covers CUR configuration, incremental ingestion, idempotency, schema drift, and billing adjustments.
I would export actual and amortized cost at the billing-scope level, then preserve subscription and resource dimensions for allocation.
- Daily exports land in immutable storage paths partitioned by billing month, export type, and run date.
- The model retains reservation, savings plan, marketplace, refund, and pricing fields instead of reducing every charge to retail cost.
- Azure hierarchy metadata maps tenant, management group, subscription, resource group, and tags to the 80-subscription ownership model.
- Close logic freezes a version on business day 3 and posts later adjustments into a controlled restatement table.
Why interviewers ask this: The interviewer is checking whether Azure exports support both engineering analysis and finance-close reconciliation.
I would use the detailed BigQuery billing export as the source of truth and snapshot hierarchy metadata separately.
- Billing exports retain project, service, SKU, resource, labels, credits, invoice month, and cost-at-list fields at their native grain.
- A daily project snapshot captures organization, folder, billing account, owner, and lifecycle because hierarchy changes over time.
- CUD fees and credits are modeled separately before allocation so effective savings and utilization are not hidden inside net cost.
- Incremental queries partition by usage date and cluster by project and service to keep 900-project processing predictable.
Why interviewers ask this: The interviewer is evaluating whether the candidate understands GCP billing grain, hierarchy history, credits, and CUD economics.
I would use immutable monthly source partitions, incremental normalization, and explicit versioned restatements instead of full refreshes.
- External-stage loads capture file checksum and provider export version so duplicate files cannot double cost.
- dbt incremental models rebuild only open months plus partitions named in a restatement control table.
- Large facts cluster on billing month, provider, billing account, and resource keys used by downstream allocation joins.
- Data tests run on staged aggregates first, while detailed exceptions are sampled and quarantined to protect the 6-hour SLO.
Why interviewers ask this: The interviewer is testing whether the candidate can balance auditability, restatement, and warehouse performance at multi-billion-row scale.
I would publish a measurable contract for completeness, validity, freshness, uniqueness, and financial reconciliation at every provider boundary.
- Completeness compares ingested account, subscription, and project totals with provider control totals before allocation.
- Validity tests required owner, cost center, service, charge category, currency, and usage dates against governed dimensions.
- Freshness pages the data owner after 24 hours, while stale dashboards display their last successful billing timestamp.
- Coverage excludes only approved unallocatable charges, and every missing 2% has a dollar amount, owner, reason, and expiry date.
Why interviewers ask this: A strong answer makes quality enforceable in dollars and ownership rather than presenting a generic list of dbt tests.
I would reconcile in layers from invoice to billing account to charge category before comparing allocated views.
- Provider invoice totals are matched by invoice month, currency, tax treatment, billing account, and legal entity.
- A bridge explains amortization, credits, refunds, marketplace charges, support, taxes, and late corrections between cash and effective cost.
- Control totals are captured before and after FOCUS normalization and again after shared-cost allocation to detect where variance entered.
- Any residual above 0.5% blocks the close dataset and receives an owned reconciliation item rather than an unlabeled adjustment.
Why interviewers ask this: The interviewer is evaluating whether the candidate can preserve accounting traceability through normalization and allocation transformations.
I would first calculate cost-weighted allocation coverage, because metadata on 72% of resources does not show how much of the $45M is assigned.
- Coverage is assigned dollars divided by total in-scope dollars, reported by charge category and cost center rather than inferred from resource counts.
- Reliable account, subscription, project, and tag mappings assign cost with source lineage and a documented confidence level.
- Shared and unknown pools remain visible with dollar value, allocation rule, and owner while each center receives a monthly showback.
- A 90-day plan prioritizes the largest unassigned dollar pools and delays chargeback until coverage, reconciliation, and disputes meet policy.
Why interviewers ask this: The interviewer is testing whether allocation coverage is measured by assigned cost, not by the percentage of resources carrying metadata.
I would require an auditable policy that defines charge scope, allocation precedence, close timing, disputes, and restatements.
- Directly attributable usage posts first, contract fees and shared services follow published drivers, and unknown cost stays in a visible suspense center.
- Each rule has an owner, effective date, source fields, calculation, materiality threshold, and approval history.
- Cost centers receive a 5-business-day preview before monthly posting and 10 business days to dispute source or rule errors.
- Retroactive changes above $50K require finance approval and a restatement, not silent rewriting of prior P&Ls.
Why interviewers ask this: A strong answer treats chargeback as a financial control with reproducibility and appeal rights, not merely a dashboard setting.
I would use causal usage drivers where available and a separately approved business driver only for truly non-metered costs.
- Kubernetes platform cost follows requested CPU and memory adjusted by shared idle, because teams can influence those quantities.
- Network, observability, and data-platform cost use bytes, telemetry volume, or compute consumption instead of headcount.
- Central security and unavoidable enterprise overhead may use revenue or headcount, but remain labeled as non-controllable.
- I would back-test each rule on 3 months of data and require finance plus engineering approval if any center shifts by more than 10%.
Why interviewers ask this: The interviewer is evaluating causal allocation, controllability, and governance when no single shared-cost driver is universally fair.
I would separate close from dispute resolution and use a materiality-based workflow with complete lineage.
- Every charge exposes source record, allocation rule version, driver quantity, rate, and owning service so challenges are evidence-based.
- Disputes below $10K post in the current month and adjust next month; larger items enter a finance-approved reserve.
- A joint FinOps, finance, and platform panel resolves policy disputes weekly, while source metadata errors route directly to data owners.
- Root causes are measured by dollars and recurrence, and a rule with repeated 5% disputes is reviewed before the next cycle.
Why interviewers ask this: The interviewer is testing whether disputes have financial controls, traceability, and service levels without holding the entire close hostage.
I would define a transaction boundary and allocate only the cost causally required to complete that transaction.
- A governed event count deduplicates retries, declines, and reversals so the denominator matches the product definition.
- Direct service cost comes from resource ownership, while shared Kubernetes, database, network, and observability cost follows measured consumption.
- The metric is segmented by product, region, payment rail, and success state because a single blended average hides mix changes.
- Monthly reconciliation ties the numerator to the $28M cost model and explains rate movement through price, efficiency, and transaction mix.
Why interviewers ask this: A strong answer protects both numerator and denominator semantics and makes unit-cost movement diagnosable.
I would model cost per successful inference by model, workload class, and token volume rather than divide total spend by call count.
- GPU time, reserved capacity, idle headroom, CPU, memory, vector search, and network cost are attributed from serving telemetry.
- Input tokens, output tokens, batch size, cache hits, latency tier, and failed calls explain why two inferences cost differently.
- Shared training and platform cost stays separate or is allocated by an approved product driver so serving efficiency remains visible.
- The product P&L receives $/1K tokens and $/successful inference with monthly reconciliation to the $15M ledger.
Why interviewers ask this: The interviewer is evaluating whether AI unit economics reflects workload drivers, capacity idle, and product accounting rather than a misleading blended ratio.
I would build a governed cost ledger that maps provider charges through services and shared platforms into product P&L lines.
- Product, service, owner, cost center, and legal-entity dimensions use effective dates so reorganizations do not rewrite history.
- Direct cost posts from cloud ownership metadata, while shared services use versioned drivers agreed with product and finance.
- A monthly bridge connects invoice, amortized cost, allocated cost, and P&L posting with residuals held in suspense.
- Product leaders receive controllable cost and unit economics separately from corporate overhead, with drill-through for all $55M.
Why interviewers ask this: The interviewer is testing dimensional modeling, effective dating, shared allocation, and accounting traceability for product profitability.
Locked questions
- 21
For $65M of predictable annual compute across AWS, Azure, and GCP, how would you design a 3-year commitment ladder while keeping 20% flexibility?
designcommitments - 22
AWS represents $30M of annual compute, with 70% steady EC2, 20% database, and 10% volatile workloads; how would you combine Savings Plans and RIs?
databaseawscommitments - 23
Azure compute costs $14M annually across 80 subscriptions, and 25% may move regions within 1 year; how would you choose between Reservations and Azure savings plans?
commitments - 24
GCP contributes $11M of annual spend across compute, GKE, and data services, with demand expected to vary by 30%; how would you structure CUD coverage?
coveragekubernetes - 25
A $48M forecast has plus or minus 20% uncertainty; what commitment risk bands and approval limits would you establish for the next 24 months?
forecastingcommitments - 26
You manage 40 commitment instruments totaling $22M with maturities spread over 18 months; what renewal calendar and decision package would you build?
commitmentsspread - 27
An automation platform may purchase up to $500K of cloud commitments per week; which guardrails would you require before allowing unattended transactions?
guardrailstransactionscommitments - 28
Procurement offers a 3-year $12M cloud agreement at a 19% discount, but engineering forecasts range from $9M to $15M; how would you make the decision?
forecastingprocurement - 29
A $50M cloud portfolio must deliver 18% verified savings in 12 months without reducing any production SLA; how would you build the optimization portfolio?
optimization - 30
One hundred Kubernetes clusters cost $12M annually, but namespace allocation covers only 65% and production must retain 30% failover headroom; what would you optimize?
finops-loopallocationkubernetes - 31
A GPU fleet of 2,000 accelerators costs $20M annually and inference p99 must remain below 300 ms; how would you reduce cost without hiding idle capacity?
capacity - 32
Stateless workloads cost $8M annually and must meet 99.95% availability; how would you raise Spot usage from 15% to 60% while targeting no more than 10% simultaneously disrupted capacity?
spotcapacity - 33
Eighty Kubernetes clusters have $3M of annual optimization potential; how would you choose between native Karpenter and CAST AI within a 90-day evaluation?
optimizationkubernetes - 34
Data transfer and storage cost $9M annually across 3 clouds, and regulated data must remain in-region for 7 years; what optimization architecture would you propose?
optimization - 35
Cloud marketplace and infrastructure licenses cost $5M annually, but production support coverage cannot fall below 24/7; how would you optimize the portfolio?
finops-loopcoverageoptimization - 36
How would you forecast a $50M annual multi-cloud portfolio 24 months ahead while keeping monthly variance below 5%?
dispersionforecasting - 37
The board needs a 3-year cloud range for a $50M run-rate business whose transaction growth could be 20% to 60%; how would you construct scenarios?
transactions - 38
For a $50M annual portfolio, what would you include in a 1-page monthly executive report that must be read in under 5 minutes?
- 39
At $50M annual spend, how would you govern cost anomalies so teams are not paged for every 5% fluctuation but material drift is owned within 24 hours?
iac - 40
Finance closes in 3 business days and budgets $52M annually; how would you integrate cloud forecast, accruals, and chargeback without waiting for final provider invoices?
forecastingallocationbudgeting - 41
An optimization backlog claims $8M of annual savings, but finance recognizes only $2M; how would you build a savings-realization system over 12 months?
backlogsystem-designoptimization - 42
You have 2 quarters to roll chargeback from 4 pilot teams to 26 cost centers covering $40M; how would you manage the change?
allocation - 43
How would you run a quarterly executive FinOps review for a $70M portfolio with the CFO, CTO, and 6 product VPs in 60 minutes?
finops - 44
Thirty-five engineering teams must contribute to a 15% unit-cost reduction, but bonuses cannot reward raw cloud-spend cuts; what incentive design would you propose?
design - 45
A showback schema used by 30 teams must move to a FOCUS-aligned version within 6 months; what policy and deprecation contract would you publish?
focusschemaallocation - 46
Six FinOps tools cost $2.4M annually and overlap in allocation, commitments, Kubernetes, and anomaly detection; how would you set a 3-year vendor strategy?
finopscommitmentskubernetes - 47
The CEO mandates a verified 25% reduction on an $80M annual cloud run-rate within 12 months, with workload volume growth of 30% and no SLA reductions; what program would you lead?
- 48
You must adopt FOCUS across AWS, Azure, and GCP in 9 months while keeping existing reports for 50 teams within 0.5% of current totals; what roadmap would you use?
focusdecision-makingroadmap - 49
An RFC must choose within 90 days between building a Snowflake and dbt FinOps platform for $900K or buying a $1.4M annual product; how would you drive the decision?
dbtfinopscost-data-pipeline - 50
Four engineers must help 18 product teams define ambiguous unit economics within 6 months, and no product may use more than 2 primary metrics; how would you mentor the group?
unit-economicsmentoringmonitoring - 51
At 09:10, AWS, Azure, and GCP spend rises 38% above a $2.4M monthly baseline, but revenue is flat; which costs do you contain first, what evidence do you preserve, and what gate lets teams resume changes within 2 hours?
- 52
Cross-region data transfer jumps 240% to $18,000 per day after a routing release 35 minutes ago; flow logs show 62% of bytes on one API. What do you reroute, what evidence identifies the payer, and what 1-hour recovery gate do you use?
api - 53
Logging spend rises 310% from $90,000 to $369,000 monthly within 6 hours of a release, and 74% comes from one debug event. What do you suppress, what evidence must remain, and what 24-hour recovery gate prevents losing audit coverage?
coveragelogging - 54
An orphaned GPU training fleet has burned $126,000 in 9 hours, pushing daily ML spend 185% over plan; utilization is 2%. What do you terminate, what evidence do you capture, and what 30-minute gate permits GPU scheduling again?
utilizationjobs - 55
EKS cost climbs 48% to $740,000 monthly over 3 hours while requests rise only 6%; Kubecost assigns 57% to unallocated. What do you contain, which Kubernetes evidence resolves ownership, and what 2-hour gate restores autoscaling?
kuberneteskubernetes-costownership - 56
The GCP billing export is missing 19% of a projected $1.8M month for 14 hours while the CFO close starts tomorrow. How do you contain reporting risk, what evidence proves the gap, and what 6-hour gate allows the close dataset to publish?
- 57
A cloud invoice arrives at $2.886M, which is $286,000 or 11% above the approved $2.6M month, and CUR differs from the invoice by 4.7% two days before payment. What do you dispute, what evidence supports it, and what 48-hour gate releases payment?
aws-cur - 58
A 9% currency move and an expired $420,000 credit make reported Azure spend look 27% above plan within one day, though usage rose 3%. How do you correct the signal, what evidence is authoritative, and what 4-hour gate ends the escalation?
escalation - 59
Demand falls 32% in 10 days, leaving a $6.2M annual commitment portfolio at 68% utilization. What do you contain, what evidence supports a portfolio change, and what 14-day recovery or exit gate do you set?
utilizationcommitments - 60
A migration completed 3 weeks early and stranded $1.4M of AWS RIs and GCP CUDs, dropping coverage benefit 41%. What do you reuse or exchange, what evidence proves eligibility, and what 7-day gate closes recovery?
migrationscoverage - 61
Compute Savings Plan utilization falls from 96% to 61% in 4 hours, exposing $22,000 per day; deployments show a 44% move to Graviton. What do you pause, what evidence confirms causality, and what 24-hour gate restores rollout?
causaldeploymentcommitments - 62
ProsperOps bought $3.8M annualized commitments after a bad demand feed overstated baseline by 26% for 18 hours. What automation do you stop, what evidence scopes the overbuy, and what 72-hour recovery gate do you require?
commitments - 63
A $9.6M commitment tranche renews in 36 hours, but the latest forecast is 18% lower and vendor cancellation closes in 6 hours. What do you renew, what evidence supports the decision, and what 30-day performance gate follows?
resiliencecommitmentsperformance - 64
AWS offers a 7% EDP rebate on $28M over 3 years, while GCP offers a 21% CUD discount but a planned migration could move 35% of spend in 14 months. What evidence decides the split, and what quarterly exit gate protects downside?
migrations - 65
A Spot policy saves $46,000 per day but a 2-hour interruption wave causes $180,000 in retries and 14% failed jobs. What do you move off Spot, what evidence prices the incident, and what 7-day gate permits expansion?
spotincidents - 66
Finance capitalized $2.7M of prepaid commitments, while FinOps amortized them differently, creating an 8% monthly P&L dispute 3 days before close. Which view do you publish, what evidence resolves it, and what close gate do you set?
finopscommitmentscost-basis - 67
Chargeback close is blocked for 18 hours because 13% of a $4.4M month is unallocated across 22 cost centers. What do you provisionally post, what evidence repairs ownership, and what 6-hour gate reopens close?
allocationownership - 68
A Terraform release drops cost-center tags from 31% of resources for 9 hours, leaving $390,000 unowned this month. What do you roll back, what evidence reconstructs tags, and what 24-hour gate restores deployment?
terraformdeploymentrollback - 69
Two products dispute $680,000 of shared platform cost, moving their margins by 12%, with quarterly close due in 48 hours. Which allocation do you hold, what evidence decides the driver, and what gate finalizes chargeback?
allocationcss - 70
After a FOCUS mapping release, duplicate Charge records inflate a $7.1M month by 16% for 5 hours. What do you quarantine, what evidence identifies duplication, and what 12-hour gate republishes reports?
focus - 71
A product hierarchy change moved 24 services between business units midmonth, reallocating $930,000 and shifting one unit's cost 19% overnight. What evidence sets effective dates, and what 2-day gate accepts the corrected chargeback?
allocation - 72
After chargeback starts, 6 teams threaten to stop deployments because their bills rise 28%, or $1.2M, with 72 hours left in the quarter. What do you pause, what evidence separates error from behavior, and what gate resumes enforcement?
allocationdeployment - 73
A pipeline bug drops about 26% of transactions from the denominator for 8 hours, making unit cost jump from $0.42 to $0.57 and triggering a $600,000 cut proposal. What do you block, what evidence repairs the metric, and what 24-hour gate reopens decisions?
monitoringci-cdtransactions - 74
A board deck overstated quarterly cloud cost by $2.3M, or 9%, and was sent 3 hours ago to 14 executives. What do you correct first, what evidence accompanies it, and what 24-hour gate restores reporting?
cloud-cost - 75
A product launch spends $1.9M in 36 hours, 64% above its $1.16M budget, while signups are 22% below plan. What do you throttle, what evidence protects growth, and what 12-hour recovery gate do you set?
budgetingthrottle - 76
A holiday forecast misses actual cloud spend by 23%, or $3.4M, after 5 days because last year's peak shifted by 2 weeks. What do you reforecast, what evidence changes seasonality, and what 48-hour gate approves capacity spend?
forecastingcapacity - 77
AI inference cost jumps 52% from $0.008 to $0.0122 per call in 4 hours after a model release, while quality improves 6%. What do you route back, what evidence prices the gain, and what 24-hour gate keeps the model?
- 78
Cost per transaction worsens 29% from $0.31 to $0.40 over 8 hours while total spend rises only $74,000 and transactions fall 18%. What do you change, what evidence separates numerator and denominator, and what 2-day gate accepts recovery?
transactions - 79
A forecast model that held 6% error now misses by 21%, or $2.1M, for 3 consecutive weeks after pricing and architecture changes. What do you retire, what evidence proves drift, and what 14-day promotion gate applies to a replacement?
forecastingarchitectureiac - 80
A late marketplace feed creates a $1.3M accrual gap, understating monthly cloud expense by 12% with 10 hours left in close. What do you accrue, what evidence supports it, and what 5-day reversal gate do you set?
- 81
The board's 15% cloud-margin improvement target is $4.8M short with 45 days left, and current actions cover only 62%. What do you stop or accelerate, what evidence ranks actions, and what weekly recovery gate do you report?
css - 82
A savings report claims $5.6M, but rightsizing and commitment benefits overlap by 37% across the same 60-day baseline. What do you retract, what evidence assigns causality, and what 7-day gate republishes savings?
optimizationcommitmentsfinops - 83
A rightsizing rollout saves $38,000 per day but raises checkout p99 46% to 1.3 seconds within 25 minutes. What do you roll back, what evidence isolates risk, and what 2-hour gate allows another canary?
deployment-strategiesoptimizationrollback - 84
FinOps wants to delete an idle environment costing $480,000 annually, but Engineering says it supports a failover test in 9 days and is 96% idle. What do you suspend, what evidence decides, and what post-test kill gate do you set?
hypothesis-testingfinops - 85
A CAST AI policy and Karpenter both consolidate the same cluster, terminating 28% of nodes in 12 minutes and causing $96,000 in lost orders. What do you disable, what evidence shows control conflict, and what 6-hour recovery gate applies?
- 86
A new GPU quota policy cuts daily spend 34%, or $210,000, but blocks 18% of priority inference jobs for 40 minutes. What do you exempt, what evidence revises policy, and what 24-hour gate restores enforcement?
- 87
A database storage-class optimization saves $72,000 monthly but increases read latency 58% and triggers $140,000 in SLA credits over 3 hours. What do you revert, what evidence proves the tier caused it, and what 4-hour gate permits retry?
optimizationlatencydatabase - 88
A storage deduplication and CDN routing change projects $1.1M annual savings, but restores fail in 7% of tests and egress rises $36,000 in 6 hours. What do you disable, what evidence isolates each change, and what 24-hour gate permits rollout?
queriestesting - 89
A vendor optimizer changes 4,600 resources in 20 minutes, cutting projected spend 17% but causing 23% service errors and $310,000 revenue loss. What authority do you revoke, what evidence scopes blast radius, and what 8-hour gate restores automation?
finops-loopoptimizationresilience - 90
A team claims $2.2M annual savings 30 days after optimization, but effective spend fell only 4% and traffic fell 11%. What claim do you suspend, what evidence validates causality, and what 14-day gate recognizes savings?
causaloptimizationvalidation - 91
The FinOps council deadlocks for 21 days over a $3.6M shared-platform allocation, and 40% of members reject both proposed drivers. What interim decision do you make, what evidence breaks the tie, and what 30-day governance gate follows?
allocationlockingfinops - 92
The CFO demands a 20% cut worth $8M in 60 days, while the CTO says any cut above 7% risks a 99.95% SLO. What do you authorize, what evidence resolves the conflict, and what weekly stop gate applies?
slo - 93
After 3 reporting corrections in 6 weeks, executives distrust a dashboard covering $42M annual spend; usage totals differ by 8%. What do you disable, what evidence rebuilds trust, and what 30-day publication gate do you set?
- 94
The CEO orders a 25% cloud reduction worth $12.5M with a 90-day deadline, but only 54% of ideas have owners and 18% threaten customer SLOs. What do you commit, what evidence ranks the plan, and what biweekly recovery gate applies?
estimationslo - 95
An auditor finds missing approvals for 16% of $11M in commitment purchases made over 9 months. What authority do you suspend, what evidence reconstructs the trail, and what 10-day gate closes the finding?
commitments - 96
Cloudability is unavailable for 7 hours during a $900,000 daily spend spike, and teams lose 63% of anomaly visibility. What fallback do you activate, what evidence remains authoritative, and what 4-hour recovery gate returns the tool?
cssfinops-platform - 97
A senior analyst's formula error moved $760,000 to the wrong business unit and overstated its cost 14% for 5 days. What do you correct, what evidence preserves accountability, and what 48-hour gate restores their publishing rights?
- 98
A mid-level engineer recommends deleting backups to save $340,000 annually, but evidence shows 8% are the only copies within a 24-hour recovery objective. What do you stop, how do you mentor with evidence, and what 30-day gate permits a revised recommendation?
mentoringbackups - 99
A portfolio of 12 FinOps initiatives costs $2.4M annually, but 5 have delivered under 20% of promised savings for 2 quarters. What do you kill, what evidence ranks exits, and what 45-day recovery gate protects the survivors?
finopspromises - 100
After a 4-hour cost-control incident causes $1.1M in revenue loss while saving $86,000, 7 remediation items compete for one team and only 40% fit this quarter. What do you prioritize, what evidence orders work, and what 30-day postmortem gate closes the incident?
prioritizationincidents