Ruby Developer interview questions
100 real questions with model answers and explanations for Senior Ruby Developer candidates.
See a Ruby Developer resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I cannot size Puma from RPS and p99 alone; I need measured mean request occupancy, CPU per request, and RSS per worker under representative load.
- If mean occupancy is 50 ms, Little's Law gives 4,000 times 0.05, or 200 concurrent request slots before headroom; p99 latency is not a substitute for that mean.
- I benchmark worker counts against the eight pod CPU and 1.5 GiB limits, then add threads only where PostgreSQL or HTTP wait improves throughput without exhausting ActiveRecord pools.
- k6 generates the traffic, while infrastructure metrics measure Puma backlog, CPU, RSS, and operating-system run queue; I keep only a configuration that meets p99 without sustained saturation.
Why interviewers ask this: The interviewer is testing whether Puma sizing starts from measured occupancy and resource limits rather than RPS and p99 alone.
I would cap each worker pool at seven instead of letting 288 web threads consume all 320 connections.
- Twelve times three times seven gives 252 web connections, leaving 68 for migrations, consoles, monitoring, and failover.
- Primary, replica, and Sidekiq pools are budgeted separately because each process and database role owns a pool.
- I require checkout wait p99 below 10 ms; if seven is insufficient, I reduce query occupancy or add PgBouncer rather than exceed the database budget.
Why interviewers ask this: The interviewer is evaluating connection arithmetic and operational headroom across process types.
I would keep identity and workflow state outside Puma so any healthy pod can serve the next request.
- Sessions use signed encrypted cookies for less than 4 KB, or opaque IDs in Redis when revocation is required.
- Uploads go directly to object storage with 10-minute presigned URLs; no request depends on a pod filesystem or memory cache.
- Orders and idempotency keys live in PostgreSQL with unique constraints, and termination drains in-flight requests for 30 seconds without sticky sessions.
Why interviewers ask this: The interviewer is checking whether the candidate removes hidden instance affinity while preserving correctness.
I would use CDN, Redis fragments, and request memoization, with a separate freshness contract for each layer.
- CDN caches public JSON for 30 seconds with 60-second stale-while-revalidate; a 95% hit rate cuts origin load to about 750 RPS.
- Redis stores assembled products for 120 seconds under versioned keys such as product:42:v17; process memory is not a shared cache.
- Request memoization removes repeated policy and association work only inside one request, while hit ratio and origin p99 are measured per layer.
Why interviewers ask this: The interviewer is testing explicit freshness and quantified load reduction at every cache layer.
I would derive keys from database versions and pair event invalidation with a 60-second TTL.
- Product fragments use cache_key_with_version, while categories use one content_version instead of touching millions of child rows.
- The update transaction increments that version and writes an outbox row, so invalidation cannot disappear after commit.
- Consumers delete Redis and CDN keys by explicit surrogate key; wildcard SCAN over 20 million keys is too slow and noisy.
Why interviewers ask this: The interviewer is evaluating transactional invalidation and a hard staleness bound at large key counts.
I would split workloads into processes with explicit concurrency instead of letting slow exports occupy payment threads.
- Payments get 24 threads and a queue-latency SLO below 2 seconds, emails get 32 with 30 seconds, and exports get 24 with 10 minutes.
- Strict priority belongs only in a small critical process because a permanently busy first queue starves later queues.
- Separate Kubernetes deployments isolate 2 GiB exports and scale from queue latency plus execution time, not queue length alone.
Why interviewers ask this: The interviewer is testing whether business latency classes become concrete Sidekiq boundaries.
I would make the database operation idempotent and use Sidekiq uniqueness only to reduce duplicate work.
- The job accepts payment_id and inserts a capture under a unique index on payment_id plus provider_operation.
- One transaction locks the payment, creates or finds the capture, and writes an outbox event; retries return the stored result.
- unique_for may suppress duplicates for 30 minutes, but the provider also receives the same idempotency key because a Redis lock is not a 24-hour correctness boundary.
Why interviewers ask this: The interviewer is checking durable idempotency across queue, database, and provider boundaries.
The target is 2,500 jobs per second, requiring about 400 concurrent slots and 100 CPU cores before headroom.
- Three million divided by 1,200 seconds is 2,500 jobs per second; 2,500 times 160 ms equals 400 in-flight jobs.
- CPU demand is 2,500 times 40 ms, or 100 CPU-seconds per second, so I provision about 130 cores with 30% headroom.
- Several processes use cores despite the GVL, while roughly four threads per process overlap I/O; load tests include Redis, database pools, and object storage.
Why interviewers ask this: The interviewer is evaluating throughput arithmetic and the difference between thread concurrency and CPU capacity.
I would match a partial composite index to the exact tenant predicate and ordering.
- For tenant_id, open status, and created_at DESC, I would test an index on tenant_id, created_at DESC where status = 'open'.
- If only 4% of orders are open, the partial index is roughly 25 times smaller than indexing every status.
- INCLUDE id and total_cents may enable index-only scans, but I accept them only after EXPLAIN ANALYZE shows low heap fetches and write cost under load.
Why interviewers ask this: The interviewer is testing index design from a real query with quantified selectivity and write cost.
I would use INCLUDE columns only if index-only scans save enough heap I/O to offset heavier writes.
- Search columns remain B-tree keys; only stable projected columns go in INCLUDE because they do not affect ordering.
- Five wide columns might grow a 40 GB index to 80 GB and double WAL for that index, so I estimate bytes rather than call coverage free.
- Frequent writes clear visibility bits, so I compare heap fetches, WAL, and replica lag with EXPLAIN (ANALYZE, BUFFERS) on production-like traffic.
Why interviewers ask this: The interviewer is evaluating visibility-map limits and write amplification from covering indexes.
I would route correctness-sensitive reads to primary for a bounded window and leave browsing on replicas.
- After order creation, the response carries a signed primary-read token valid for 10 seconds; checkout GETs with it use the writing role.
- Search and reviews tolerate 500 ms lag and retain most of the 70% offload.
- ActiveRecord connected_to boundaries are explicit because HTTP-verb switching misses POST-then-GET workflows; I alert before replica lag reaches the token window.
Why interviewers ask this: The interviewer is testing explicit consistency semantics instead of treating all replica reads alike.
I would range-partition by month first and shard only after one primary reaches a measured limit.
- Monthly partitions are about 2 TB, prune time-bounded queries, and make 24-month retention a fast DROP instead of a huge DELETE.
- Queries include occurred_at, and local indexes lead with tenant_id, occurred_at so pruning and tenant access align.
- Sharding adds routing, resharding, and cross-shard aggregation; I trigger it only above thresholds such as 70% storage or WAL beyond replica capacity.
Why interviewers ask this: The interviewer is evaluating the least complex scaling boundary tied to measured node limits.
I would enforce depth and cardinality-aware complexity before execution, then add runtime deadlines.
- I would start with max_depth 12 and max_complexity 200, charging connection fields by requested page size rather than one flat point.
- first is capped at 100 and unbounded lists are rejected, preventing aliases from multiplying million-row resolvers.
- Each request gets a 400 ms execution deadline and PostgreSQL statement_timeout near 350 ms, while persisted trusted operations receive reviewed budgets.
Why interviewers ask this: The interviewer is testing layered GraphQL controls that reflect result cardinality.
I would batch associations by key with GraphQL::Dataloader so query count stays constant as the page grows.
- One source loads customers by ID and another loads items by order_id, reducing about 201 queries to three.
- Sources preserve result order and return nil for missing keys, preventing records from attaching to the wrong parent.
- Orders cap at 100 and items at 50, while ActiveSupport instrumentation and a test fail this operation above five SQL queries.
Why interviewers ask this: The interviewer is evaluating correct bounded batching protected by query-count tests.
I would use opaque keyset cursors over a stable compound ordering rather than OFFSET.
- The cursor encodes created_at and id, matching an index on account_id, created_at DESC, id DESC.
- The next page uses (created_at, id) < (?, ?) and LIMIT 101 to derive has_next_page without COUNT(*).
- Page size is capped at 100 and cursors are signed; OFFSET 1,000,000 discards a million entries and drifts under inserts.
Why interviewers ask this: The interviewer is testing stable pagination, matching indexes, and API encapsulation at scale.
I would make additive changes in place and reserve a new version for semantic breaks with a measured migration window.
- OpenAPI is the reviewed source; CI blocks removed fields, narrower enums, and newly required inputs.
- Consumer contracts cover the five highest-volume clients, while telemetry tracks version and field use across all 40.
- A breaking v2 runs beside v1 for 90 days with Sunset metadata and an owner for every remaining client.
Why interviewers ask this: The interviewer is evaluating enforceable compatibility and a concrete retirement process.
I would acknowledge only durable receipt, then dispatch work from the database inbox rather than from an after_commit callback.
- A unique index on provider and provider_event_id defines event identity, while an HMAC covers the raw body plus a timestamp with a five-minute replay window; duplicates receive the same 2xx.
- The controller inserts the inbox row in one short transaction and returns 202 within 300 ms, with a 10 MB body limit, schema version, and 72-hour retention in the contract.
- Pollers claim committed inbox rows with FOR UPDATE SKIP LOCKED, enqueue the stored row ID, and mark dispatch; a crash can duplicate enqueueing, so the worker processes that ID idempotently.
Why interviewers ask this: The interviewer is testing durable webhook acceptance without an after-commit enqueue loss window.
I would keep one deployment but create executable package boundaries around business domains.
- Packwerk packages own models, jobs, controllers, and public APIs; CI rejects new privacy and dependency violations.
- Billing reaches Accounts through a narrow interface, while direct access to Billing tables and private constants is forbidden.
- Each package has an owner and focused tests, but can share one PostgreSQL transaction; existing violations get a baseline and domain-by-domain burn-down.
Why interviewers ask this: The interviewer is evaluating enforceable modularity without premature distributed-system costs.
I would extract it only when the domain and operational boundary outweigh distributed transactions and extra operations.
- A dedicated team, separate compliance cadence, stable API, and distinct scaling profile are stronger evidence than class count.
- I first isolate Billing behind an in-process interface and collect six weeks of dependency telemetry while stopping cross-domain table writes.
- The remote API uses coarse idempotent commands; extraction is rejected if checkout still needs one synchronous Orders and Billing transaction.
Why interviewers ask this: The interviewer is testing measurable extraction evidence rather than generic microservice arguments.
I would insert the order and outbox row in the same PostgreSQL transaction, then publish asynchronously.
- Each row carries event_id, aggregate_id, type, schema_version, payload, and created_at; event_id is the consumer idempotency key.
- Relays claim batches with FOR UPDATE SKIP LOCKED, publish to Kafka, and mark sent, accepting possible duplicate publication after a crash.
- At 20,000 per second I partition by day and alert when oldest unpublished age exceeds 10 seconds; consumers enforce idempotency.
Why interviewers ask this: The interviewer is evaluating atomic event creation and honest at-least-once delivery at scale.
Locked questions
- 21
A Kafka Rails pipeline handles 50,000 account events per second across 120 partitions and requires per-account ordering; how would you choose keys?
kafkarailspartitioning - 22
A Rails SaaS serves 30,000 tenants, including 20 regulated tenants requiring strict isolation, on 2 TB of PostgreSQL; what model would you choose?
postgrescloudrails - 23
One tenant generates 35% of traffic in a 10,000-tenant SaaS, while every tenant has a 500 ms p99 SLA; how would you control the noisy neighbor?
cloud - 24
You must add required country_code to a 600-million-row users table while Rails writes 8,000 rows per second; what is the zero-downtime plan?
rails - 25
A 900-million-row payments table must change integer cents to a richer money representation with less than 0.01% mismatch; how would you migrate it?
- 26
A Rails endpoint spends 70 ms parsing JSON on CPU and 30 ms waiting on PostgreSQL at a 10,000 RPS target; what does the GVL imply?
postgresendpointsrails - 27
A pod has 4 vCPU and 3 GiB, each warmed Ruby worker uses 620 MiB, and workload is 60% I/O; would you favor processes or threads?
concurrencyruby - 28
A Ruby service runs 200 CPU-heavy calculations per second at 25 ms each within 2 GiB; would you use threads, processes, or Ractors?
concurrencyruby - 29
Rails pods limited to 2 GiB grow from 900 MiB to 1.8 GiB RSS in six hours; how would you set a Ruby GC and memory budget?
railsmemoryruby - 30
A preloaded Rails pod runs four Puma workers under a 3 GiB cgroup limit, and each worker reports 600 MiB RSS; how would you preserve copy-on-write and budget memory?
railsmemory - 31
A Rails API spends 65% of CPU in Ruby at 7,000 RPS and has 250 MiB memory headroom per pod; how would you evaluate YJIT?
decision-makingapimemory - 32
A Rails search endpoint must move from 900 ms to 250 ms p99 at 1,500 RPS; what profiling plan would you use?
railsendpointsprofiling - 33
A Rails codebase defines 400 fields through method_missing and must stay below 100 ms p99; what metaprogramming constraints would you impose?
metaprogrammingrails - 34
A 700,000-line Rails application with 25 engineers has no static typing; how would you adopt Sorbet without blocking 20 weekly deploys?
decision-makingdeploymentrails - 35
A typed payment service consumes 12 external JSON schemas and must keep invalid payloads below 0.001%; where does Sorbet end?
schema - 36
A Rails admin API updates 60 account attributes for 5,000 support agents; how would you prevent authorization and mass assignment flaws?
authapirails - 37
A Rails application serves 2 million browser sessions and embeds user HTML; what controls would you set for XSS and CSRF under 99.99% availability?
xsscsrfsessions - 38
A Rails import service fetches 100,000 customer URLs daily from Kubernetes; how would you constrain SSRF without blocking HTTPS imports?
railshttpskubernetes - 39
A Rails checkout has 99.95% monthly availability and 400 ms p99 at 9,000 RPS; what observability would you build?
railsobservability - 40
A Rails request crosses GraphQL, PostgreSQL, Redis, Sidekiq, and Kafka, but only 1% of traffic may be traced; how would you preserve causality?
kafkarailsbackground-jobs - 41
A Rails deployment runs 40 Kubernetes pods at 6,000 RPS and must keep 99.99% availability during releases; how would you configure rollout?
kubernetesdeploymentconfig - 42
A Puma pod gets 30 seconds to terminate while requests can run for 20 seconds; how would you configure probes and draining without 502s?
config - 43
Traffic rises from 1,000 to 12,000 RPS in five minutes while Sidekiq latency must stay below 10 seconds; how would you autoscale both?
scalingbackground-jobslatency - 44
A stateless Rails API serves US and EU with 180 ms network latency and a 500 ms checkout p99 SLA; where would you place state?
railsapilatency - 45
A public Rails API allows 600 requests per minute per account, bursts to 100, and sustains 20,000 RPS across 100 pods; how would you rate-limit it?
railsapi - 46
A Rails reservation endpoint handles 1,200 RPS for 50 shared seats, each reservation has an event_id, and only one active reservation may hold a seat; what PostgreSQL transaction would you use?
transactionspostgresendpoints - 47
Eighty Rails and Sidekiq processes need up to 20 connections each, but PostgreSQL permits 400; how would PgBouncer change the design?
designrailsbackground-jobs - 48
A 1.2-million-line Rails platform must upgrade from Rails 7.0 to 8.x while maintaining 50 daily deploys; how would you structure it?
railsdeployment - 49
A checkout workflow spans order, inventory, payment, and notifications and must finish synchronously within 700 ms; how would you shape Rails domain services?
rails - 50
A deployment includes a 400-million-row index build and must roll back within 10 minutes while serving 5,000 writes per second; how would you release it?
deploymentrollbackindexes - 51
A Rails checkout endpoint's p99 rose from 280 ms to 1.9 s immediately after a deploy, while p50 stayed near 90 ms and errors remain below 0.2%. What do you do in the first 30 minutes?
railsendpointsdeployment - 52
Each Puma worker starts at 420 MB RSS and grows to 1.4 GB over eight hours, but drops only to 1.2 GB after a full GC. Kubernetes kills two workers daily. How do you distinguish retained Ruby objects from allocator fragmentation?
kubernetesruby - 53
A Rails process has 6 million more String objects after importing 500,000 CSV rows, and you can route only 2% of traffic away from it for ten minutes. How would you collect and use heap evidence safely?
concurrencydata-structuresrails - 54
During a 4,000 requests-per-second sale, Rails p99 shows 350 ms spikes every 20 seconds and GC.stat shows major_gc_count increasing at the same moments. CPU is 65%. What is your response?
rails - 55
After deploying a native image-processing gem, eight 4-vCPU Puma pods show only 28% total CPU, yet requests serialize at 2.4 seconds and adding threads has no effect. How do you handle the incident?
soft-skillsincidentsdeployment - 56
Puma runs 6 workers with 12 threads each. At 900 requests per second, CPU is 45%, request queue time reaches 2.5 seconds, and thread dumps show most threads waiting on one 4-second partner API. What do you change first?
concurrencydata-structuresapi - 57
A Rails fleet has 10 pods, each with 2 Puma workers and 10 threads, but PostgreSQL allows 160 application connections. During bursts, ActiveRecord checkout waits exceed 3 seconds. What decision do you make?
postgresormconcurrency - 58
A Rails orders page loads 50 orders in 1.4 seconds and emits 151 SQL queries: one for orders, 50 for customers, and 100 for line items and products. How do you fix and verify it?
sqlqueriesrails - 59
After replacing preload with eager_load, an ActiveRecord report returns 1.8 million joined rows for 12,000 accounts because it joins invoices and support tickets, both has_many associations. RSS reaches 2.2 GB. What do you do?
joinsorm - 60
The default Sidekiq queue grows from 5,000 to 240,000 jobs in 20 minutes, oldest-job latency is 17 minutes, Redis is healthy, and workers use only 35% CPU. How do you respond?
redislatencydata-structures - 61
A Redis primary failover loses 84,000 queued Sidekiq jobs because persistence was disabled, although producers already returned 202. How do you recover and change the acceptance contract?
redisdata-structuresbackground-jobs - 62
A partner API returns 503 for 40 minutes. Sidekiq retries grow to 1.2 million jobs, Redis memory reaches 88%, and recovery traffic would exceed the partner's 300 requests-per-second limit. What do you do?
redisapimemory - 63
The Redis cluster used for Rails.cache is unavailable for 12 minutes. Database CPU jumps from 40% to 94%, checkout waits hit 1.8 seconds, and Sidekiq uses a separate healthy Redis. How do you keep the API alive?
databaseredisrails - 64
A cached Rails home feed takes 900 ms to build. Its five-minute key expires at once across 30 pods, creating 6,000 database queries in ten seconds. What change do you ship?
databasequeriescaching - 65
A Rails migration adding an index has run for nine minutes on a 180-million-row table. PostgreSQL shows it blocking inserts, checkout waits exceed five seconds, and the deploy is still progressing. What do you do?
indexespostgresmigrations - 66
Two Rails jobs update Account and Subscription rows in opposite order. PostgreSQL reports 800 deadlocks per hour, and Sidekiq retries complete eventually but triple load. How do you fix it?
consistencyrailsbackground-jobs - 67
A planned PostgreSQL promotion finishes in 25 seconds, but 40% of Rails writes fail for six minutes because pooled connections still point to the old primary. How do you recover and prevent a repeat?
postgresrails - 68
A Rails endpoint verifies 200 signatures per request in pure Ruby. One Puma worker pins a core at 100%, adding threads does not improve throughput, and p95 reaches 1.6 seconds. What do you change?
concurrencyrubyendpoints - 69
A Rails singleton memoizes exchange rates with @rates ||= fetch_rates. Under 16 Puma threads, tests show two simultaneous refreshes and occasionally a partially updated Hash. How do you make it safe?
concurrencymemoizationtesting - 70
A prototype moves image metadata parsing to four Ractors, but a production gem stores configuration in mutable class state and raises Ractor::IsolationError. The prototype is 1.7 times faster. Do you keep Ractors?
configgemsprototypes - 71
A Rails serializer canary receives 5% of traffic, and 1.3% of requests routed to it return 500 because legacy rows have nil currency while the new code assumes a String. What do you do?
serializationdeployment-strategiesrails - 72
A Rails report leaves a PostgreSQL transaction idle for nine hours; 180 million dead tuples cannot be removed, the orders table grows fourfold, and p99 reaches 3 seconds. What do you do?
transactionspostgresrails - 73
Rotating Rails signing keys logs out 68% of two million browser sessions and invalidates password-reset links during the release. What do you do in the incident and next rotation?
passwordssessionsincidents - 74
After a GraphQL persisted-query allowlist becomes mandatory, 22% of requests from older mobile builds fail with PersistedQueryNotFound. How do you recover without reopening arbitrary queries?
queriesgraphql - 75
A Rails order update commits in PostgreSQL and then publishes directly to Kafka. A network error after the commit leaves 2,300 shipped orders without events, so the warehouse is stale. How do you repair and redesign it?
postgreswarehousekafka - 76
A recommendations service slows from 80 ms to 6 seconds. Rails has a 7-second read timeout and retries once, so Puma threads double their outbound calls and the checkout API begins timing out. What do you change during the incident?
resiliencerailsapi - 77
Mobile clients retry POST /orders after a 2-second timeout. Rails creates 186 duplicate orders in one hour because the first requests committed but their responses were lost. How do you make the endpoint idempotent?
resiliencerailsendpoints - 78
A payment provider sends the same webhook up to 12 times and may deliver refund before charge.succeeded. Your Sidekiq handlers currently append one ledger row per delivery. What is your design?
designbackground-jobswebhooks - 79
A Kafka producer ships amount as an object instead of integer cents; one Rails consumer partition retries the poison record for 25 minutes and lag reaches three million. What do you do?
kafkarailspartitioning - 80
One malformed Sidekiq job raises JSON::ParserError in 40 ms, retries thousands of times across workers, and blocks a queue whose SLO is two minutes. How do you handle it?
slobackground-jobsdata-structures - 81
A Rails backfill must update 70 million rows. The first version uses Model.all.each inside one transaction, reaches 9 GB RSS, and holds locks for 50 minutes. How do you rewrite it?
backfillrailstransactions - 82
A Rails request has a 1.5-second SLO and calls inventory and tax services sequentially. Their p99 values are 700 ms and 900 ms, and each client retries once with a 1-second timeout. What do you change?
resilienceslorails - 83
After adding an HTTP call inside an ActiveRecord transaction, database pool usage stays at 100% whenever the partner takes three seconds. The pool has 20 connections and Puma has 20 threads. What do you refactor?
databasetransactionsorm - 84
After preload, four Puma workers report RSS values of 480 MB, 490 MB, 910 MB, and 930 MB. The two large workers receive most long-lived WebSocket connections and restart every six hours. How do you investigate and mitigate?
websockets - 85
A Rails JSON API's heap_live_slots stay flat after warm-up, but RSS rises from 500 MB to 850 MB during bursts and settles at 760 MB. Switching to jemalloc lowers steady RSS to 610 MB but adds a new runtime dependency. What do you decide?
railsapidependencies - 86
stackprof shows a Rails serializer allocating 1.6 million Arrays per second because it repeatedly calls map, compact, and flatten over the same records. CPU is 78% and GC consumes 22% of request time. What do you change?
serializationrails - 87
Puma begins raising Errno::EMFILE after a traffic spike. Each pod has a 1,024 file-descriptor limit, 400 open client sockets, and 500 outbound sockets stuck in keep-alive pools. What do you do?
problem-solvingdescriptors - 88
A custom Ruby thread pool runs work outside Rails controllers. One in 20 jobs sees the previous tenant in Current.account, creating a cross-tenant data leak. How do you fix it?
concurrencyrubyrails - 89
A deploy renames InvoiceGenerator to Billing::InvoiceGenerator while 420,000 queued Sidekiq payloads still reference the old class; NameError retries saturate Redis. What do you do?
redisdeploymentdata-structures - 90
A scheduled Rails job must run once per customer across 40 Sidekiq processes. A Redis lock has a 60-second TTL, but some jobs take 90 seconds, so two workers overlap and overwrite results. What do you change?
railsbackground-jobsredis - 91
Product wants a coupon feature in two weeks. The same Rails pricing module causes 14 incidents per quarter and has 2,800 lines of callbacks; a safe extraction is estimated at six weeks. What do you recommend?
estimationincidentscallbacks - 92
A Ruby engineer submits a 900-line PR that adds a fourth ActiveRecord callback to process refunds and has only controller specs. The release is in three days. How do you review and mentor them?
concurrencycallbacksruby - 93
A code review adds rescue StandardError around a Sidekiq invoice job and returns success to stop alerts. Last month 600 invoices were silently skipped. What feedback and replacement do you give?
feedbackcode-reviewalerting - 94
A junior engineer's Rails change triggers 8% checkout errors after deploy. They are on call for the first time and want to debug before rollback, while revenue is dropping by $4,000 per minute. What do you do?
deploymentrollbackrails - 95
Your current Rails version reaches end of security support in four months. The upgrade has 320 deprecation warnings, 18 internal gems, and a product launch in six weeks. How do you sequence the work?
rails - 96
The Rails test suite takes 48 minutes and fails randomly in 7% of CI runs, mostly time-zone and Sidekiq specs. The team reruns CI until green and loses about 30 engineer-hours weekly. What do you change?
railsbackground-jobstesting - 97
Rails p99 breaches its 800 ms SLO twice a week, but logs contain no request ID, ActiveRecord pool waits are not measured, and traces sample only 0.1%. You have five engineering days. What do you prioritize?
slorailsprioritization - 98
At 02:00, payment success drops from 98.8% to 81%, Sidekiq payment lag is nine minutes, and three recent deploys touched checkout, Redis configuration, and a provider client. How do you lead the incident?
deploymentconfigredis - 99
A senior Ruby engineer repeatedly approves queries that pass tests but scan 40 million rows in production. Their latest PR adds where.not on an unindexed JSONB field to a request path expected at 600 requests per second. How do you handle the review?
soft-skillsqueriestesting - 100
3 incidents in 6 weeks came from Sidekiq jobs being non-idempotent, yet the roadmap has no reliability work and another bulk-email feature starts Monday. What decision do you take to the team?
incidentsprioritizationidempotency