Skip to content

QA Engineer interview questions

100 real questions with model answers and explanations for Senior candidates.

See a QA Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

e2epyramidcomponents

Use a service-local pyramid with 70% unit, 18% component, 7% integration, 4% contract, and 1% E2E tests, then weight the upper layers toward the 8 critical services.

  • For a 10,000-test portfolio, allocate 7,000 unit, 1,800 component, 700 integration, 400 contract, and 100 E2E tests; run all unit and component tests plus changed-service contracts in pull requests.
  • Give each critical service 10 E2E journeys and 20 provider or consumer contracts, use the remaining 20 E2E journeys and 240 contracts for the other 42 services, and avoid duplicating assertions between layers.
  • Shard pull-request tests into 24 workers with a 10-minute target and reserve 2 minutes for setup; run cross-service integration and the full 100-journey E2E suite nightly within 90 minutes.
  • Gate merges on 100% pass for changed-service unit, component, and contract suites, p95 pull-request duration below 12 minutes, and nightly E2E pass rate of at least 99% over 30 runs.

Why interviewers ask this: This split keeps most feedback cheap and local while concentrating expensive cross-service evidence on business-critical paths.

code-reviewshardingmonorepo

Run affected unit, component, and contract tests on every pull request, while moving broad integration and E2E coverage to merge queues and scheduled suites.

  • Cap each pull request at 16 compute-minutes: 2 unit shards, 2 component shards, and 1 contract shard cost 15 compute-minutes and can execute concurrently in 4 minutes.
  • At 120 pull requests, this consumes 1,800 compute-minutes, leaving 600 for ten 12-minute integration shards and twenty-four 20-minute E2E shards distributed across merge queues and nightly runs.
  • Select affected tests from the dependency graph, but always include contracts for every changed API or event schema and one smoke E2E journey for changes touching checkout or authentication.
  • Gate merges on all selected tests passing, wall time below 10 minutes, and affected-test selection recalling at least 99% of tests that failed in a 30-day full-suite replay.

Why interviewers ask this: The schedule satisfies both the compute budget and fast feedback SLA without removing expensive system-level checks.

coveragecapacityconcurrency

Allocate the 300 scenarios in proportion to impact multiplied by change frequency, with a minimum smoke floor for every service.

  • The weighted totals are 240 for payment services, 270 for customer-workflow services, and 60 for internal services, for 570 risk points overall.
  • Reserve 60 scenarios as one smoke per service, then distribute 240 scenarios by risk: 101 to payments, 114 to customer workflows, and 25 to internal services.
  • The final allocation is 113 payment scenarios, 132 customer-workflow scenarios, and 55 internal scenarios, emphasizing contracts and integration paths at high-risk boundaries.
  • Gate completion on 100% coverage of impact-5 requirements, at least one negative test per critical boundary, and no service with fewer than one executable smoke scenario.

Why interviewers ask this: A transparent risk score directs scarce scenarios to high-impact, frequently changing behavior while retaining baseline coverage everywhere.

coveragecode-review

Gate changed code immediately and raise repository-wide branch and mutation thresholds in two measured releases rather than treating line coverage as sufficient.

  • On pull requests require at least 85% line and 75% branch coverage for changed lines, no repository-wide regression above 0.5 percentage points, and execute mutants only in changed modules within the 3-minute cap.
  • In release one set repository gates at 82% line, 65% branch, and 50% mutation; in release two raise them to 84%, 70%, and 60% after surviving mutants are classified and tests are added.
  • Nightly, sample high-risk modules until the 25-minute budget is exhausted, prioritizing payment, authorization, and parsing code; rotate lower-risk modules so every service receives mutation analysis within 14 days.
  • Reject a change when changed-code mutation score is below 70%, more than 5 surviving or timed-out mutants remain unclassified, or branch coverage falls below the release gate.

Why interviewers ask this: Branch and mutation gates expose weak assertions that a high line percentage can conceal while respecting both CI limits.

pricingtesting

Combine 30 deterministic examples with 10,000 generated cases per currency around explicit pricing invariants and boundary-biased generators.

  • Keep examples for 0%, 1%, 69%, and 70% discounts, quantities 1 and 10,000, and tax rates 0% and 30%, including currency-specific rounding rules.
  • Generate 120,000 cases total using valid weighted ranges, with 40% of values drawn from boundaries and adjacent values, and shrink every failure to the smallest currency, quantity, discount, and tax tuple.
  • Assert nonnegative totals, monotonic totals when quantity increases, discount never increasing the pre-tax amount, and round-trip agreement with the decimal reference implementation within one minor currency unit.
  • Gate merges on all 30 examples and all seeded generated cases passing in under 6 minutes, with each invariant exercised at least 10,000 times and every failure seed persisted as a regression example.

Why interviewers ask this: Properties cover a vast numeric space economically, while curated boundaries preserve readable evidence for known business rules.

designtesting

Encode the workflow as a finite-state model and generate transition-covering paths with role and invariant checks instead of enumerating every sequence.

  • Build model guards for all 18 legal and 11 forbidden transitions, then use a greedy path set that covers every legal transition for each allowed role and every forbidden transition at least once.
  • Generate 500 seeded walks of length 1 to 12, biasing 50% of choices toward rarely covered transitions and shrinking a failure to the shortest reproducing transition sequence.
  • Compare model state with API state after every step and assert terminal-state immutability, role authorization, idempotency for repeated commands, and conservation of order total.
  • Gate merges on 100% transition coverage, 100% forbidden-transition rejection, all 500 walks completing within 8 minutes, and no model-to-system state divergence.

Why interviewers ask this: A finite-state oracle gives systematic sequence coverage without paying the combinatorial cost of all possible paths.

observability

Test event schemas and handlers locally, verify broker semantics in integration, and reserve E2E tests for a small set of timed business flows under realistic load.

  • Add consumer-driven contracts for all 14 topics, including schema compatibility, required headers, partition keys, and examples for duplicate and out-of-order events; run them on every producer and consumer change.
  • Component-test each handler with duplicates, gaps, reordering, and retry exhaustion, asserting idempotent state and exactly one business effect despite at-least-once delivery.
  • Run broker-backed integration tests for 9 services at 2,000 events per second for 10 minutes, measuring consumer lag, dead-letter routing, correlation completeness, and convergence time.
  • Gate on 100% contract compatibility, zero duplicate business effects, p99 convergence below 30 seconds, estimated consumer lag time recovering below 5 seconds within 2 minutes, and no uncorrelated events.

Why interviewers ask this: Layered event testing separates deterministic handler correctness from broker behavior and end-to-end convergence guarantees.

coverageconfigfeature-flags

Use a constrained pairwise array for routine coverage and add mandatory high-risk combinations outside the generated set.

  • Generate pairwise rows across 8 flags, payment method, locale, and device while excluding configurations forbidden by flag dependencies; target at most 80 valid rows.
  • Add 24 fixed rows covering every payment method in the 4 locales on both device classes, then add 16 rows for all-on, all-off, and risk-ranked flag interactions, for a maximum of 120 configurations.
  • Execute component tests for all 120 rows and run E2E only for the 24 payment-locale-device rows, sharded so wall time remains below 20 minutes.
  • Gate merges on 100% valid pair coverage, every mandatory row passing, zero uncovered payment-locale-device triples, and generated configuration count not exceeding 120.

Why interviewers ask this: Constrained pairwise selection cuts the configuration space while explicit rows preserve coverage of combinations whose business risk exceeds pair strength.

flakyapi

Automate deterministic checks with stable oracles, narrow browser automation to critical journeys, and keep subjective or high-cost checks outside every-change regression.

  • Automate all 320 API checks, which consume about 1.8 compute-hours per full run, and require changed-area execution on pull requests plus one full run nightly.
  • Retain 30 critical browser journeys, costing 1.5 compute-hours per run, and quarantine any test above a 1% flaky-run rate until its synchronization or isolation defect is fixed.
  • Keep the 50 visual checks as 10 automated structural assertions plus a 40-item weekly human review sample, and run each of the 40 hardware checks twice weekly for $640.
  • Gate merges on 100% pass for selected API and critical browser checks, less than 1% browser flakiness over 100 runs, weekly spend at or below $1,000, and total execution below 30 compute-hours.

Why interviewers ask this: Automation is limited by oracle quality, execution stability, and marginal cost rather than by the raw number of available checks.

exploratory

Allocate exploratory time by current change risk, reserve a discovery buffer, and stop each charter on evidence rather than elapsed time alone.

  • Assign 16 hours to each of the 6 critical domains for 96 hours, 6 hours to 7 of the 14 monthly medium-risk domains for 42 hours, and rotate 6 stable domains at 2 hours each for 12 hours.
  • Reserve the remaining 10 hours for cross-domain charters selected from automation gaps, new flag combinations, and data-boundary risks discovered during the cycle.
  • Use 2-hour charter blocks with a named risk, data set, and oracle: 90 minutes for exploration and 30 minutes for setup, coverage notes, and evidence capture.
  • Gate release readiness on all 80 planned sessions completed, every critical charter reaching its stated coverage boundary, and no unresolved high-severity finding in the tested scope.

Why interviewers ask this: A quantified charter budget preserves focused human discovery while making scope, evidence, and completion measurable.

seleniumarchitecturegrid

Use Playwright for new TypeScript coverage and retain Selenium only as a bounded legacy lane while migrating valuable Java scenarios.

  • Playwright provides one TypeScript API for Chromium, Firefox, and WebKit, while Cypress does not provide equivalent WebKit coverage and Selenium remains appropriate for the existing Java investment.
  • The initial split is 75% Playwright and 25% Selenium, avoiding a risky rewrite of 1,200 tests while preventing further growth of the Java lane.
  • Migrate by business flow rather than line by line, deleting a Selenium scenario only after its Playwright replacement passes against the same browser and environment.
  • Accept the architecture after a 200-test pilot reaches at least 98% pass consistency across 20 runs, stays within 15% of the runtime estimate, and exposes no required browser capability gap.

Why interviewers ask this: The split satisfies the concrete browser and language constraints without turning historical framework investment into permanent architectural duplication.

playwrightsharding

Start with 12 shards at 4 workers each, then tune from measured utilization rather than increasing parallelism blindly.

  • Serial work is 6,000 × 24 seconds = 40 hours, while 48 concurrent workers give a 50-minute theoretical floor.
  • At 70% parallel efficiency, the expected runtime is about 71 minutes, leaving 19 minutes for shard imbalance, setup, artifact upload, and retries.
  • Use Playwright sharding to distribute test files across 12 CI jobs and workers to execute within each job, with isolated accounts, ports, and output directories.
  • Accept the configuration when p95 runtime over 20 builds is below 90 minutes, the slowest shard is no more than 20% slower than the median, and CPU or memory saturation stays below 85%.

Why interviewers ask this: Separating shard-level distribution from worker-level concurrency makes the capacity calculation observable and protects the runtime target from resource contention.

seleniumgridjobs

Provision about 100 browser slots, including headroom, and allocate them from production browser usage rather than evenly.

  • Required concurrency is 8,000 × 30 ÷ 4,500 ÷ 0.65 = 82.1, so 83 slots meet the mathematical minimum and 100 slots add about 20% operational headroom.
  • If usage is 60% Chrome, 25% Firefox, and 15% Edge, reserve 60, 25, and 15 slots respectively while allowing idle capacity to spill over through Grid stereotypes and queue rules.
  • Size each node from measured CPU and memory per browser process, cap sessions per node, and recycle nodes to contain leaked processes instead of advertising unsupported capacity.
  • Accept the Grid when p95 queue time is below 2 minutes, slot utilization remains under 85% for the normal peak, and 20 consecutive runs finish within 75 minutes without node-related retries above 0.2%.

Why interviewers ask this: The slot count combines workload arithmetic with measured Grid efficiency and enough headroom to absorb real browser process variance.

cypressoauth

Keep Cypress for the 2,930 compatible tests temporarily and implement the 270 browser-bound scenarios in Playwright behind one shared CI contract.

  • Cypress cy.origin can cover many cross-origin transitions, but its command model does not treat 2 simultaneous tabs as first-class controllable pages, while Playwright browser contexts and Page objects do.
  • The exception lane is 270 ÷ 3,200 = 8.4% of the suite, which is small enough for a controlled introduction but large enough that rewriting product behavior to satisfy Cypress would be the wrong trade-off.
  • Share environment provisioning, test data APIs, result reporting, and failure taxonomy across both runners, but do not share framework-specific page abstractions.
  • Accept the split after all 270 scenarios pass 20 consecutive runs with under 0.5% flake and the second runner adds less than 10 minutes or 15% to CI cost; otherwise plan a full Playwright migration.

Why interviewers ask this: A measured two-runner boundary preserves the existing investment while using a browser model that can represent the required cross-origin and multi-tab behavior.

domdecision-making

Establish a versioned selector contract owned jointly by product and QA, with semantic locators first and explicit test IDs for stable business controls.

  • Prefer role, label, and accessible name when they are part of the user contract; use data-testid for dynamic grids, duplicated labels, canvas controls, and elements whose copy is localized.
  • Ban CSS classes, DOM-depth selectors, and positional nth-child selectors in application tests because styling and layout are not automation contracts.
  • Convert the 500 highest-change tests first; reducing their 40 weekly selector failures by 80% removes 32 failures per week before scaling the pattern to all 3,000 tests.
  • Accept the contract when at least 95% of selectors pass static lint rules, selector-caused failures fall below 0.2% of executions for 4 weeks, and every new test ID has a named product owner.

Why interviewers ask this: A selector contract makes testability an explicit interface and measures success through failure reduction rather than raw test ID count.

componentsvalidation

Map the suite to 3,380 component checks, 1,300 page checks, and 520 flow checks, with dependencies allowed only from flows to pages to components.

  • Component objects expose widget behavior and state, such as selecting a date or reading validation, without knowing routes, users, or upstream workflows.
  • Page objects compose components and own page-level readiness and actions, while flow objects coordinate business journeys across pages without duplicating selectors.
  • Keep assertions in tests except reusable state predicates, and avoid a single base page that accumulates unrelated waits and navigation behavior.
  • Accept the layers when at least 90% of selectors exist in one component or page owner, no flow contains raw selectors, and the 520 journey tests run without duplicated login or navigation sequences.

Why interviewers ask this: The three layers align abstraction boundaries with the measured test distribution and prevent giant page objects from becoming a second application framework.

fixturesplaywrightauth

Create role-specific storage state in a setup project and give each parallel worker isolated mutable data through worker-scoped fixtures.

  • Repeating login costs 3,600 × 7 seconds = 7 hours of aggregate browser time, while loading storage state in about 1 second reduces that work to roughly 1 hour before parallelization.
  • Generate 8 role states through supported API or UI setup, store them as short-lived CI artifacts, and never reuse a state beyond its token lifetime or target environment.
  • Use worker-scoped account leases for tests that mutate server data and test-scoped fixtures for pages and disposable records, with cleanup through APIs rather than slow UI teardown.
  • Accept the design when authentication setup is below 5% of total runtime, 50 repeated parallel runs show zero cross-test data collisions, and expired-state recovery succeeds without manual reruns.

Why interviewers ask this: Separating reusable identity state from isolated mutable test data removes repeated login cost without introducing shared-account coupling.

testing

Treat 252 flaky outcomes per day as the absolute budget and drive the operating target below it with deterministic synchronization and accountable quarantine.

  • Daily volume is 7,000 × 12 = 84,000 executions, and 84,000 × 0.003 = 252 allowed flaky outcomes, measured before retries so retries cannot hide instability.
  • Replace fixed sleeps with Playwright web-first assertions, locator actionability, explicit response predicates, and application readiness signals tied to the state the test needs.
  • Allow 1 CI retry only for diagnosis, attach trace, network, console, and screenshot artifacts, and quarantine a test only with an owner, defect link, and 7-day expiry.
  • Accept the controls when pre-retry flake stays below 0.3% for 4 weeks, fixed waits are zero by static scan, and quarantined tests remain below 1% of the 7,000-test inventory.

Why interviewers ask this: Counting flakes before retries and requiring deterministic readiness signals turns the budget into an engineering control rather than a reporting adjustment.

coveragevisualdesign

Use a risk-tiered Playwright screenshot strategy with 720 full-page baselines for 60 critical routes and component-level baselines for the remaining visual states.

  • The naive matrix is 420 × 3 × 2 × 2 = 5,040 snapshots, while 60 × 3 × 2 × 2 = 720 protects critical revenue and navigation routes with reviewable scope.
  • Cover the other 360 routes through stable component stories and targeted element screenshots, avoiding duplicate full-page images where shared shells dominate the pixels.
  • Pin browser versions, fonts, locale, clock, animations, data, and viewport; mask only genuinely nondeterministic regions and use a small documented pixel threshold rather than broad tolerance.
  • Accept the system when every baseline change has human approval, unexplained diffs block merge, false-positive runs stay below 2%, and the critical visual job finishes within 20 minutes.

Why interviewers ask this: Risk-tiering preserves meaningful visual signal while controlling the combinatorial baseline and review burden.

code-review

Use a usage-weighted matrix with a 6,000-case pull request gate, a broader nightly gate, and a small real-device contract suite.

  • The full matrix is 5,000 × 4 × 3 = 60,000 cases, so run all 5,000 tests on the primary Chromium desktop profile plus 500 critical tests on WebKit iPhone and 500 on Firefox desktop per pull request.
  • Nightly, run the 500 critical tests across 8 priority browser-device combinations for 4,000 cases and keep the complete 60,000 cross-product for a weekly capacity window.
  • Run 100 authentication, upload, payment, camera, and download checks on real Safari iOS and Chrome Android devices because Playwright emulation does not reproduce every OS browser behavior.
  • Accept the matrix when it covers at least 95% of production traffic by browser-device weight, pull request p95 stays below 30 minutes, and no severity 1 compatibility defect escapes for 2 release cycles.

Why interviewers ask this: The schedule concentrates fast feedback on the dominant platform while retaining periodic breadth and real-device evidence for capabilities that emulation cannot guarantee.

Locked questions

  • 21

    An API must sustain 600 RPS for 10 minutes with p95 at most 400 ms, p99 at most 800 ms, and errors below 1%. A baseline k6 iteration takes 1.2 seconds. How would you configure and accept an open arrival-rate test?

    load-testingapiconfig
  • 22

    A distributed JMeter test must generate 12,000 RPS. One generator sustains 1,800 RPS at 65% CPU, average responses are 8 KB, and you require 25% generator headroom. How many generators do you provision and how do you validate the topology?

    validationdistributedgenerators
  • 23

    You need 15,000 HTTP RPS plus 500 persistent WebSocket sessions in a 10-minute CI job; 4 of 6 engineers know JavaScript, 1 knows Python, 1 knows Java, and each generator is limited to 1 GB RAM. Would you choose k6, JMeter, Gatling, or Locust, and what evidence must support the choice?

    load-testingjavascriptgenerators
  • 24

    A release candidate delivers 2,400 RPS for 15 minutes with p95 280 ms, p99 760 ms, and 0.2% errors. The SLO at 2,200 RPS is p95 at most 300 ms, p99 at most 700 ms, and errors below 0.5%. Do you approve it?

    slo
  • 25

    Production traffic is 70% hot cached reads, 20% warm-key reads that miss once per test, and 10% unique cold reads. How would you build a 5,000 RPS, 10-minute test without creating an unrealistically high cache hit rate?

    caching
  • 26

    Normal load is 1,000 RPS, a campaign can jump to 5,000 RPS for 60 seconds, and the capacity limit is unknown. What spike and stress tests would you run, and what numeric decision would each produce?

    performance-testingcapacity
  • 27

    Design an 18-hour soak at 1,500 RPS for a service starting at 4.2 GB memory, with p95 300 ms and 400 open connections after warm-up. Which slope gates would make the result pass or fail?

    designmemory
  • 28

    A closed test uses 200 users and reports p95 900 ms at 200 RPS, but the requirement is 1,000 RPS and response time can rise to 2 seconds. Is the result valid, and how would you remove coordinated omission?

  • 29

    Three generators process 1 million, 3 million, and 6 million requests and report p95 values of 100 ms, 200 ms, and 400 ms. Can you approve a 300 ms p95 SLO by averaging those percentiles?

    concurrencygeneratorsslo
  • 30

    A staging environment has one quarter of production compute and sustains 450 RPS at p95 250 ms; production must support 1,500 RPS, while pull-request CI allows 5 minutes and scheduled CI allows 30 minutes. How do you use the small result and design two performance tiers?

    designperformance
  • 31

    A pull-request pipeline handles 120 changes per day: building the artifact takes 2 minutes, 6,000 unit tests take 3 minutes, 600 selected component tests take 6 minutes, contract tests take 4 minutes, and 30 browser smokes take 8 minutes after a 2-minute deploy. How would you arrange the DAG to keep p95 below 13 minutes and avoid wasted work when 18% of changes fail?

    unitcontractdeployment
  • 32

    A regression suite consumes 960 runner-minutes, the CI platform allows 100 workers, and the required p95 wall time is 12 minutes. How would you design duration-balanced shards?

    shardingdesignregression
  • 33

    At peak, 40 approved pull requests enter a merge queue each hour, the validation suite takes 15 minutes, only four isolated environments are available, and 4 percent of changes fail. How would you keep queue wait p95 below 35 minutes?

    code-reviewvalidationdata-structures
  • 34

    A monorepo has 200 services and 25,000 tests; the full suite takes 180 minutes, but pull-request feedback must stay under 15 minutes. How would you use dependency-graph selection without losing the full safety net?

    monorepodependenciesfeedback
  • 35

    You create 180 ephemeral test environments per day; each currently needs 12 minutes to start, runs tests for 35 minutes, remains idle for 60 minutes, and costs $0.02 per environment-minute. How would you meet a $250 daily budget and a 7-minute startup objective?

    test-environments
  • 36

    A service is built 70 times per day, each build takes 8 minutes, and rebuilding between pull-request, staging, and release tests has produced a 0.8 percent binary mismatch rate. How would you design immutable artifact testing?

    designimmutabilityartifacts
  • 37

    A pipeline runs 10,000 tests with a historical 99.6 percent first-pass rate, 82 percent line coverage, and a 20-minute target. What numeric quality gates would you define so aggregate percentages cannot hide a bad change?

    coveragequality-gatesaggregation
  • 38

    Forty integration shards each need an application, PostgreSQL, Redis, and Kafka; the Kubernetes test cluster allows 60 pods, each shard takes 6 minutes, and the pipeline budget is 18 minutes. How would you arrange the dependencies?

    postgresshardingredis
  • 39

    A monorepo runs 12,000 tests on 120 pull requests per day; 70% of tests are hermetic, each averages 30 seconds, and CI must cut runner-minutes by 50% without reusing a result after any relevant input changes. How would you design test-result caching?

    code-reviewdesignmonorepo
  • 40

    Your CI executes 50,000 tests per day and emits 1,500 failures representing about 120 root causes, but engineers need the correct owner and evidence within 15 minutes. How would you design result observability and routing?

    designobservabilitytesting
  • 41

    Three services expose 48 REST operations, and CI has room for 240 contract tests, but hand-written mocks regularly drift from the OpenAPI files. How would you build a measurable OpenAPI test architecture?

    contractmockingrest
  • 42

    A tenant API has 12 protected endpoints, 3 roles, 3 credential states (valid, expired, missing), and 2 targets (own tenant, foreign tenant). How would you implement the 12 × 3 × 3 × 2 = 216 auth matrix plus payload-negative coverage without masking authorization defects?

    coveragedefectsendpoints
  • 43

    One provider is consumed by 14 applications in an estate of 18 services, and each consumer can have 3 active versions. How would you use Pact Broker and can-i-deploy to prevent an incompatible provider or consumer release?

    contractdeployment
  • 44

    A REST v1 API has 27 active clients while v2 is rolling out, and the platform emits 8 Kafka event types whose consumers may lag by 2 releases with 14 days of retained traffic. How would you test backward compatibility for both synchronous and asynchronous contracts?

    restkafkaasync
  • 45

    A suite has 600 tests over 15 tables, factory data misses a production segment with 18% null profiles, production holds 4.2 million rows, and raw production copies are forbidden. How would you combine factories with anonymized production-shaped data?

    anonymizationfundamentalstesting
  • 46

    You must mask 2.4 million records across 6 related tables in 50,000-row chunks while preserving joins, uniqueness, formats, and identical output on reruns. How would you implement and verify deterministic masking?

    joins
  • 47

    CI runs 60 times per day with 8 parallel workers, shared schemas and queues collide, and a performance test needs an exclusive 45-minute window with a stable 10-million-record baseline. How would you divide ephemeral and shared environments and guarantee cleanup?

    schemaperformancedata-structures
  • 48

    For one release, pre-release testing found 120 valid defects: S1=2, S2=8, S3=30, S4=80. During the next 30 days, production found 20 attributable defects: S1=1, S2=3, S3=6, S4=10. Using severity weights 8, 5, 2, 1, calculate escape rates and define the release gate.

    severity-prioritydefectsaws
  • 49

    Ten release defects have detected-to-final-production-verification durations of 2, 3, 4, 5, 6, 8, 10, 12, 20, and 40 hours, and the 40-hour defect was reopened after an initial closure. How would you calculate defect MTTR percentiles and capture the lifecycle correctly?

    percentilesclosuresdefects
  • 50

    A 30-day quality dataset shows 120 tests that failed then passed unchanged out of 20,000 test executions, time-to-signal p50=7 and p90=18 minutes, 6 rollback or defect-hotfix changes out of 80 production changes, weighted escape=11%, and 4,960 passed contract checks out of 5,000. Build a balanced dashboard with exact decision gates.

    defectsrollback
  • 51

    In the last 200 CI runs, 28 were red, 24 of those passed on rerun, and the team enabled up to 3 blind retries with 4 hours left before release. What do you do?

  • 52

    A Playwright payment callback scenario fails in 11 of 80 runs because the UI sometimes updates in 2.1 seconds, while the test uses a fixed 2-second wait. The release window closes in 90 minutes. What do you do?

    playwrightcallbacks
  • 53

    Twelve API tests pass alone but 9 fail when the 240-test suite runs in random order; all 9 reuse one customer account and leave different subscription states. What do you do?

    api
  • 54

    The pull-request suite grew from 18 to 57 minutes in 6 weeks, the CI limit is 35 minutes, and 14 of the last 40 builds were canceled before results arrived. What do you do?

  • 55

    An 8-shard live CI run finishes in 34 minutes because one shard takes 31 minutes while the fastest takes 7; 46 heavy specs were assigned by file count only. What do you do?

    sharding
  • 56

    After a frontend release, 168 of 420 Playwright tests fail because generated CSS classes changed, but a manual check of 25 critical actions finds no product defect. What do you do?

    defectsplaywrightcss
  • 57

    A checkout release changed a 3-step flow to 2 steps, and 96 of 150 Cypress scenarios now fail on the removed confirmation page with 70 minutes left in the release validation window. What do you do?

    validationcypress
  • 58

    An automatic Chrome update causes 73% of 600 Selenium tests to fail with session startup errors; Firefox remains green and the release deadline is in 2 hours. What do you do?

    seleniumsessionsestimation
  • 59

    The test artifact service returns 503 for 100% of jobs, 62 CI runs are blocked during the last 90 minutes, and a release decision is due in 60 minutes. What do you do?

    artifacts
  • 60

    Forty-five minutes before release, 9 of 1,800 tests fail: 7 match known flaky signatures, while 2 new failures affect password reset and invoice totals. The business asks for a blanket waiver. What do you do?

    flakypasswords
  • 61

    Release 8.14 shows the wrong renewal date to 18,400 users; 620 renewals failed and support received 73 tickets before rollback. What do you do in the first 24 hours, and how do you prevent the same escape?

    rollback
  • 62

    A payment webhook retry in release 5.6 charged 214 customers twice across 286 transactions, while all 340 automated payment tests passed. How do you respond and change the test system?

    system-designresiliencewebhooks
  • 63

    Production incident INC-417 locked 3,200 users out after password reset, although authentication coverage was reported as 91%. What evidence do you gather, and which 1 or 2 fixes do you choose?

    authpasswordscoverage
  • 64

    The defect escape rate rose from 1.8% to 3.1%, 4.7%, and 6.2% over releases 21 through 24, while release frequency stayed at 2 per week. How do you investigate and reverse the trend?

    defects
  • 65

    A critical payout flow was green in 126 tests because the bank dependency was mocked, but production rejected 9% of 4,800 payouts after the bank changed its timeout behavior. What do you change?

    resiliencedependenciestesting
  • 66

    Release 3.9 passed staging but failed for 41% of 12,000 production uploads because staging used object storage version 7.2 and production used 6.8. How do you recover and close the parity gap?

  • 67

    After API v32, 17% of 90,000 mobile profile loads failed because the service returned null for a field declared as a required string. Why did existing tests miss it, and what do you implement?

    apifundamentals
  • 68

    Search p95 latency increased from 480 ms to 2.4 seconds in release 11.3, affecting 1.6 million queries before detection, while the performance suite remained green. How do you diagnose and prevent a repeat?

    latencyqueries
  • 69

    An invoice export corrupted 6,700 of 240,000 records containing non-Latin names and negative adjustments, although 58 export tests passed. How do you reconcile the data and remove the blind spot?

    testing
  • 70

    Release 6.2 made the checkout button invisible at 200% zoom and removed its accessible name, affecting an estimated 14,000 keyboard and low-vision sessions before 96 reports arrived. What do you do next?

    estimationsessions
  • 71

    A load test sustained 12,000 requests per second at 180 ms p99, but production reached 2.4 seconds p99 at only 3,000 requests per second. The test used 90% cached reads, while production traffic contains 45% writes and a few hot accounts. What do you decide and change?

    cachingload-testing
  • 72

    At 8,000 virtual users, the load injector reaches 100% CPU and 95% network utilization while the service stays at 42% CPU; reported throughput plateaus and response times look stable. How do you handle this result?

    soft-skillsthroughput
  • 73

    A performance gate fails builds above 450 ms p95, but 10 unchanged-baseline runs range from 410 to 470 ms and 3 of them fail. A candidate build measures 455 ms once. What corrective decision do you make?

    performance
  • 74

    All 38 Pact consumer contracts pass, yet a provider release causes HTTP 400 responses for 17% of checkout calls because a field allowed by the matcher is null in production. What do you do beyond rerunning Pact?

    contracthttpfundamentals
  • 75

    A payment provider sandbox accepts API v2 requests in 300 ms, but production has moved to v3, rejects 12% of them, and delivers webhooks after 8 minutes instead of 10 seconds. How do you respond?

    webhooks
  • 76

    Six teams share one test environment; the full suite is 14% flaky, and failures correlate with parallel runs that reuse account ids and leave 30,000 records behind. What is your immediate and lasting decision?

    flakytest-environments
  • 77

    Ephemeral test environments should expire after 6 hours, but inventory shows 312 active environments, 47 older than 72 hours, and a monthly cost increase from $4,000 to $18,000. What action do you take?

    test-environments
  • 78

    A test with 10 million uniformly distributed rows reports 90 ms query p99, but production is 1.8 seconds because the busiest 1% of tenants own 64% of rows. What do you change before approving the optimization?

    queriesdistributedoptimization
  • 79

    A migration passed on 2 million staging rows, but production has 240 million rows with 7% null legacy values; the backfill holds locks for 11 minutes and queue lag reaches 25 minutes. What do you decide?

    fundamentalsdata-structuresmigrations
  • 80

    A release has 12 feature flags, but testing covered only all-off and all-on states; the combination of flags A and D, EU region, and client versions below 5.4 causes duplicate payments for 9% of that cohort. How do you respond?

    cohortsfeature-flagstesting
  • 81

    Three hours before release, 27 of 1,800 automated tests fail across 6 clusters while the test environment reports elevated database errors. How do you separate product defects, flakes, and environment failures and make the release decision?

    defectstest-environmentsdatabase
  • 82

    A known defect causes saved filters to disappear for an estimated 7% of 120,000 monthly users, but no data is permanently lost and a workaround takes about 20 seconds. Would you release in 4 hours?

    estimationdefects
  • 83

    The planned test window is cut from 18 hours to 6 hours after a build delay, with 420 regression cases and 4 testers available. How do you preserve a defensible release decision?

  • 84

    Ninety minutes before launch, an executive asks you to bypass a 98% critical-suite gate because the current result is 92%, with 12 of 150 tests failing. What do you do?

    testing
  • 85

    A hotfix changes 46 lines in payment retry logic during an incident affecting 18% of transactions, and the team wants production deployment in 35 minutes. How do you judge the risk?

    incidentstransactionsdeployment
  • 86

    For a go or no-go review, automation is 99.3% green across 2,400 tests and a 5% canary shows no metric regression, but exploratory testing found 2 cases of draft data loss in 40 attempts. What is your decision?

    exploratorymonitoringdeployment-strategies
  • 87

    Two hours before release, the forward deployment passes all 310 critical tests, but nobody has tested rollback for a migration touching 14 million rows. Can the release proceed?

    deploymentrollbackmigrations
  • 88

    One day before release, the compatibility matrix is complete except Safari 16 on iOS 15, which represents 11% of active users, and the changed upload flow handles files up to 25 MB. What do you decide?

  • 89

    A security scan reports 73 findings 5 hours before release; 72 match a noisy generated-file rule, but 1 is a new critical finding rated 9.8 that may allow unauthenticated account access. How do you handle the gate?

    soft-skills
  • 90

    After releasing a new search ranking service, how would you verify it in production while limiting risk when normal traffic is 60,000 requests per minute and rollback takes 4 minutes?

    rollback
  • 91

    A junior owns a suite of 120 UI tests: 38 use fixed sleeps of 2 to 5 seconds, 64 use absolute XPath selectors, and 18 fail in at least 7 of 100 CI runs. What should the senior do, how should the junior correct and present the work, and what numeric evidence should be required before approval?

    locators
  • 92

    A mid-level QA sets every flaky test to retry 3 times: the dashboard now shows a 99.5% final pass rate, but only 91% pass on the first attempt across 500 runs and 46 tests need a retry. What should the senior do, how should the mentee repair and present the approach, and which numeric exit criteria should apply?

    flakyresilience
  • 93

    A mentee submits 75 API tests that assert only a 2xx status and a nonempty body; mutation testing leaves 23 response-contract or side-effect defects undetected. What should the senior do, how should the mentee strengthen and present the suite, and what numeric verification should gate acceptance?

    defectsapi
  • 94

    A mentee models a load test as 1,000 users looping without think time, although production peaks at 900 requests per second with a 60% browse, 25% search, 10% update, and 5% export mix. What should the senior do, how should the mentee correct and present the model, and which numeric checks should validate it?

    validationload-testing
  • 95

    Eighty parallel tests share 5 customer accounts and one mutable order, causing data collisions in 14% of runs. What should the senior do, how should the mentee redesign and present test-data ownership, and what numeric verification should prove isolation?

    ownershiptesting
  • 96

    A mentee opens a 4,800-line automation pull request across 37 files that changes the runner, retry policy, reporting, and fixtures but adds no tests for the framework itself. What should the senior do, how should the mentee restructure and present the change, and what numeric gates should apply?

    fixturescode-reviewresilience
  • 97

    A mentee files a bug titled Checkout broken with one screenshot, no build or environment, and no logs, although the failure occurs in 3 of 10 attempts. What should the senior do, how should the mentee improve and present the report, and what numeric evidence should make it triage-ready?

  • 98

    A mentee's launch test plan contains 210 happy-path cases for a payment service handling 25,000 transactions per day but misses schema rollback, duplicate webhooks, and clock-skew risks. What should the senior do, how should the mentee correct and present the plan, and which numeric criteria should show adequate risk coverage?

    transactionsschemacoverage
  • 99

    A mentee's dashboard claims 94% automation coverage by counting executed cases against requirements, but it includes skipped and stale tests while only 52% of 40 critical user paths have a current automated check. What should the senior do, how should the mentee repair and present the metric, and what numeric validation should prevent another misleading report?

    coveragemonitoringvalidation
  • 100

    A production incident creates duplicate refunds on 1,800 orders and exposes $72,000 because a mentee omitted an idempotency test from the release suite. What should the senior do, how should the mentee correct and present the gap, and which numeric verification should close the action?

    incidentsidempotency