System Administrator interview questions
100 real questions with model answers and explanations for Senior candidates.
See a System Administrator resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would make each hall an independent failure domain and reserve N+1 capacity outside it.
- Spread every 3-node service across 3 halls, with separate power, top-of-rack switches, and storage paths per RHEL host.
- Size 2 surviving halls for 100% load at 70% CPU, so normal placement cannot exceed 46% per hall.
- Accept only designs whose dependency map and quarterly isolation test show 1 hall can disappear without shared DNS or identity loss.
Why interviewers ask this: The interviewer is checking whether redundancy covers correlated dependencies and includes enough surviving capacity.
I would use interchangeable Ubuntu instances behind 2 independent HAProxy pairs and externalize all durable state.
- Keep sessions in Redis and files in object storage, then require any 1 of 4 zones to lose 100% capacity safely.
- At 250 requests/s per host and 60% target utilization, provision 800 hosts because 120,000/(250×0.60)=800.
- Drain connections for 30 seconds and accept a release only when 5% canaries hold p99 below 200 ms with 25% headroom.
Why interviewers ask this: A strong answer combines statelessness, failure tolerance, capacity math, and measurable rollout gates.
I would run 1 active VRRP owner and 1 standby, with priority changes tied to HAProxy readiness.
- Give node A priority 150 and node B 140, use unicast VRRP, and track HAProxy with a 20-point penalty.
- Set advertisements to 500 ms, disable preemption after failback, and send 3 gratuitous ARP bursts when the VIP moves.
- Keep 50% spare connection capacity per node and accept the design only if 20 failover trials restore traffic within 3 seconds.
Why interviewers ask this: The interviewer expects correct VRRP ownership behavior without unsafe failback or undersized nodes.
I would require 3 of 5 Corosync votes and successful STONITH before Pacemaker starts the resource elsewhere.
- Place 2 nodes in room A, 2 in B, and 1 in C, using 2 independent Corosync networks.
- Configure IPMI fencing with 2 power paths and a 60-second timeout; quorum alone cannot stop an isolated writer.
- Allow service only with at least 3 votes, and accept fencing after 20 tests prove the old node loses power before promotion.
Why interviewers ask this: The answer must distinguish majority authority from the fencing that prevents concurrent writers.
I would choose active-passive ownership because 12 ms synchronous cross-site writes would directly tax all 8,000 writes/s.
- Keep 1 primary site and 1 warm standby, replicate asynchronously, and cap observed lag at the 30-second RPO.
- Fence the old primary through storage and power before promotion, accepting unused standby compute to avoid 2 writable authorities.
- Size each site for 100% load at 65% CPU and require quarterly switchover to meet the 10-minute RTO.
Why interviewers ask this: The interviewer is evaluating consistency, latency, fencing, and the cost of passive capacity.
I would separate authoritative zones from recursive resolution and deploy both across all 4 sites.
- Run 2 Unbound resolvers per site with anycast or DHCP distribution, keeping 8 caches independent of authoritative BIND servers.
- Use 30-second TTLs only for movable service records and 3,600 seconds for stable hosts, trading query load for convergence selectively.
- Accept the design when 1 entire site can fail and 99.9% of uncached lookups still finish within 100 ms.
Why interviewers ask this: A senior design separates DNS roles and chooses TTLs from measured convergence and load requirements.
I would use a 5-server Consul control plane and DNS names that hide ephemeral node addresses.
- Spread 5 Consul servers across 3 failure domains, require 3 votes, and keep agents off the voting path.
- Register readiness checks every 10 seconds with a 30-second deregistration threshold, accepting brief staleness instead of rapid endpoint flapping.
- Cache results for 5 seconds and accept the design at 1,000 changes/minute with p99 lookup below 50 ms after 1 voter fails.
Why interviewers ask this: The interviewer is checking quorum placement, health semantics, cache convergence, and control-plane capacity.
I would rebuild versioned Packer images and replace hosts rather than patching 2,000 long-lived instances in place.
- Pin the RHEL base digest and repository snapshot, sign each image, and retain 2 tested versions for rollback.
- Promote through 1%, 10%, and 25% rings, requiring 30 minutes of health at each ring before fleet replacement.
- Accept an image only after OpenSCAP, boot, systemd, and application tests pass, with 10% spare capacity for rolling replacement within 48 hours.
Why interviewers ask this: The answer should show reproducible artifacts, staged replacement, security evidence, and enough rollout capacity.
I would create 36 remote state roots aligned to environment and lifecycle boundaries, not 1 global state.
- Store all 36 states in a versioned encrypted backend with locking, and grant each CI identity access to only 1 state root.
- Run refresh-only plans every 6 hours; reconcile drift through code unless a 1-hour break-glass record explicitly authorizes it.
- Accept partitioning when a failed apply locks at most 1 domain and cross-state consumers use fewer than 10 stable outputs.
Why interviewers ask this: The interviewer is checking that state boundaries limit privilege and blast radius without creating uncontrolled coupling.
I would derive dynamic inventory from the CMDB and execute pinned Ansible environments through regional controllers.
- Cache inventory for 5 minutes, group by region and service metadata, and abort if targeting differs by more than 2% from CMDB counts.
- Use 5 regional controllers with 50 forks each, keeping WAN traffic local and limiting any play to 1 service group.
- Accept a playbook only after check mode and idempotency tests show the 2nd run changes 0 resources on 100 representative hosts.
Why interviewers ask this: A strong answer makes inventory authoritative, execution bounded, and idempotency objectively testable.
I would use serial push waves for the risky change and reserve pull mode for routine low-risk convergence.
- Start with 18 hosts, then 180, then 20% waves, setting max_fail_percentage to 2 and pausing 15 minutes.
- Use the Ansible sysctl module to update running and persistent state without a reboot, and require the 2nd execution to report 0 changes.
- Stop expansion if health checks exceed 1% failures; push gives wave control, while 30-minute pull convergence reduces central-controller load later.
Why interviewers ask this: The interviewer wants a deliberate push-versus-pull choice, bounded waves, and real idempotency gates.
I would promote signed repository snapshots through 4 rings so every host installs the same tested package set.
- Mirror RHEL and Ubuntu upstreams daily, retain 3 immutable snapshots, and reject packages whose GPG signatures or SBOM checks fail.
- Patch 1%, 10%, 30%, then 59% of hosts with 24-hour observation between rings, completing within 14 days.
- Accept each ring below 0.5% service failures and keep 15% capacity spare, trading slower exposure reduction for controlled blast radius.
Why interviewers ask this: The answer should connect repository provenance, deterministic versions, rollout math, and explicit safety thresholds.
I would model hard requirements separately from startup order and avoid making 900 services depend on transient network reachability.
- Add RequiresMountsFor=/srv/data plus Wants= and After=network-online.target, but let the application retry Vault for 120 seconds instead of ordering on it.
- Use Restart=on-failure, RestartSec=5, and StartLimitBurst=6 per 60 seconds to prevent a fleet-wide restart storm.
- Accept the unit after 50 boot tests prove data is mounted before start and shutdown completes within the 30-second TimeoutStopSec.
Why interviewers ask this: The interviewer is checking precise systemd dependency semantics and bounded behavior around external services.
I would place at least 2 writable Windows Server domain controllers in each office and make sites topology-aware.
- Run DNS and Global Catalog on all 6 controllers, map 3 AD Sites to subnets, and avoid a remote authentication dependency.
- Keep all 5 FSMO roles assigned deliberately across 2 central controllers; roles are movable ownership, not active-active redundancy.
- Accept the design when loss of 1 office leaves logon p95 below 2 seconds and replication converges within 15 minutes.
Why interviewers ask this: A senior answer distinguishes domain-controller redundancy, site locality, DNS, and FSMO role ownership.
I would design for 200,000 samples/s because 2,000×1,500/15 equals 200,000 before headroom.
- Provision 2 identical Prometheus replicas per regional shard independently and size each for 300,000 samples/s, adding 50% ingestion headroom.
- Keep 15 days locally, remote-write to Thanos, and alert when WAL growth or ingestion exceeds 70% of tested capacity.
- Accept the platform after a 24-hour load test sustains 300,000 samples/s with query p99 below 2 seconds.
Why interviewers ask this: The interviewer is evaluating sample-rate math, HA duplication, retention design, and measurable headroom.
I would shard collection by region, duplicate each shard, and use Thanos for deduplicated 30-day global queries.
- Deploy 2 Prometheus replicas in each of 5 regions with external labels, so 10 collectors survive 1 replica loss.
- Send blocks every 2 hours to object storage and retain 24 hours locally, accepting remote-query latency for older data.
- Accept regional isolation when 1 WAN link fails, local alerts still evaluate within 30 seconds, and global query p95 remains below 5 seconds.
Why interviewers ask this: The design must preserve local monitoring while providing durable, deduplicated global visibility.
I would impose label budgets before Grafana dashboards and route only actionable alerts through Alertmanager.
- Limit each exporter to 2,000 active series and reject labels such as user_id that can create more than 10,000 values.
- Route severity 1 to PagerDuty within 60 seconds and severity 2 to team queues, grouped by service for 5 minutes.
- Accept a rule only with 1 owner, 1 runbook, and less than 5% false pages across a 30-day review.
Why interviewers ask this: The interviewer expects concrete controls for cardinality cost and alert actionability rather than more dashboards.
I would buffer logs at the edge, index selected fields centrally, and tier 240 TB of 30-day data.
- Run Vector on 4,000 hosts with 10 GB disk buffers, protecting applications when OpenSearch is unavailable for 2 hours.
- Keep 7 days on SSD-backed OpenSearch and 23 days in object storage, accepting slower searches for older records.
- Cap indexed fields at 50 per source and accept the design when 16 TB/day replay keeps ingestion below 70% capacity.
Why interviewers ask this: A strong design balances buffering, searchable retention, indexing cost, and replay capacity.
I would encode CIS Level 1 as OpenSCAP policy and keep SELinux enforcing with 3 application-specific domains.
- Apply 95% of the CIS profile automatically, documenting time-limited exceptions for the remaining 5% with 90-day expiry.
- Generate policy from audited denials, map 3 ports to app_port_t, label /srv/app app_var_lib_t, and reject broad allow rules or permissive domains.
- Accept a host at 98% scan compliance with 0 unexpected AVC denials across a 24-hour workload test.
Why interviewers ask this: The interviewer is checking measurable baseline enforcement and narrow SELinux policy rather than disabling controls.
I would deploy 2 enforcing AppArmor profiles that grant each daemon only its required files, capabilities, and network family.
- Allow daemon A /srv/a/** r, 1 writable log path, network inet stream, and capability net_bind_service; nftables exposes only TCP 443.
- Run profiles in complain mode for 48 hours, review denials, then switch 10%, 50%, and 100% rings to enforce.
- Accept enforcement with 0 unexpected denials and p99 latency growth below 2% at 20,000 requests/s.
Why interviewers ask this: A strong answer uses concrete AppArmor path rules and measured promotion rather than a generic security claim.
Locked questions
- 21
How would you design auditd collection for 3,500 Linux hosts while limiting audit overhead to 3% CPU?
designlinux - 22
Design SSH and PAM access for 1,800 Linux hosts where 60 administrators need privileged sessions within 5 minutes.
sshlinuxremote-access - 23
How would you provide secrets to 2,200 Linux services with 24-hour rotation and no static credentials in images?
secretslinux - 24
Design network segmentation for 5,000 servers in 4 trust zones while permitting 300 documented service flows.
design - 25
How would you design LVM for a 40 TB database host needing 2-hour maintenance snapshots and 20% growth headroom?
designstoragedatabase - 26
Choose RAID for 12 disks of 8 TB each when the workload needs 70 TB usable and any 2-disk failure tolerance.
storage - 27
Design a filesystem for 300 million 8 KB files requiring 25,000 random IOPS and 30 TB usable capacity.
designcapacitystorage - 28
How would you design NFS for 1,200 Linux clients needing 12 GB/s aggregate reads and 99.95% availability?
designavailabilityaggregation - 29
Design distributed storage for 5 PB raw capacity across 60 Linux nodes while tolerating loss of 1 entire rack.
distributeddesigncapacity - 30
How would you implement 3-2-1 backup for 800 servers containing 600 TB with a 24-hour RPO?
backupsdisaster-recoverybackup - 31
Size a backup path that must transfer 24 TB during an 8-hour window while retaining 25% throughput headroom.
backupsbackupthroughput - 32
Design backup validation for 200 critical servers with a 15-minute RPO and 2-hour RTO.
designbackupsdisaster-recovery - 33
How would you design warm-standby DR for 1,000 services across 2 regions with a 5-minute RPO and 30-minute RTO?
designdisaster-recovery - 34
Design quorum and replication for a 2-site service that needs 0 split brain and can add 1 witness location.
designdistributed-systemsreplication - 35
Plan CPU capacity for 2,400 hosts with 16 cores, 45% peak use, 30% growth, and loss of 1 of 3 zones.
capacitycapacity-planning - 36
Plan RAM for 1,000 virtualization guests averaging 12 GB with 20% growth and a 1-host maintenance reserve per 20-host cluster.
virtualization - 37
Size storage for 600 TB of data growing 4 TB/week for 26 weeks with 2 total copies and 20% free space.
- 38
Design a VMware cluster for 600 VMs averaging 4 vCPUs across 12 hosts with N+2 tolerance and a 4:1 overcommit limit.
design - 39
How would you design KVM virtualization for 300 VMs on 10 hosts with 25 Gb/s networking and 1-host maintenance tolerance?
designvirtualization - 40
Design Docker isolation for 200 containers on 20 Linux hosts with a 70% CPU utilization target.
dockercontainerslinux - 41
How would you secure 1,000 Docker workloads that expose 80 services and require 10 writable paths?
docker - 42
Design a small Kubernetes worker layer for 120 pods across 6 nodes with 1-node failure tolerance.
designkubernetes - 43
How would you provide Kubernetes storage for 40 stateful pods needing 2 TB each, a 15-minute RPO, and a 2-hour RTO?
disaster-recoverykubernetes - 44
Design Layer 4 load balancing for 500 TCP services carrying 200,000 concurrent connections across 3 zones.
designload-balancingnetworking - 45
How would you design routed management networks for 3,000 hosts across 6 sites with 2 independent WAN paths?
designrouting - 46
Design time synchronization for 5,000 Linux and Windows hosts requiring clocks within 50 ms across 5 sites.
designlinux - 47
How would you design centralized Linux identity for 2,700 hosts and 8,000 users with 1-site failure tolerance?
designlinux - 48
Design lifecycle management for 500 FreeBSD and 2,500 Linux hosts with 3 supported OS generations.
designlinux - 49
How would you manage configuration for 3,200 mostly Linux hosts and 400 intermittently connected Windows servers?
configlinux - 50
Design a 3,000-host platform combining HA, patching, monitoring, backup, and security with a 99.95% target.
designbackupsbackup - 51
A 32-core RHEL application host jumps from load 18 to 180, while CPU utilization is only 35% and 12 services time out. How do you diagnose and stabilize it?
- 52
A 128 GB Ubuntu database host has 96% RAM use, 42 GB of swap occupied, kswapd at 70% CPU, and query p99 rising from 80 ms to 4 seconds. What do you do?
databasequeriesmemory - 53
At 09:20, authentication starts failing on 700 of 900 servers, but CPU and memory on the application hosts remain normal. How do you find the shared failure?
authmemory - 54
A faulty unit starts 25,000 processes on a 64-core host, PID usage reaches 99%, and SSH can barely fork. How do you regain control?
sshremote-accessconcurrency - 55
Thirty web nodes show load 60 on 16 cores, 55% iowait, and a shared storage latency jump from 4 ms to 220 ms. What is your incident decision?
incidentsincident-managementlatency - 56
A Redis dependency recovers after 8 minutes, then 400 application instances retry at once and drive it back to 100% CPU. How do you break the recovery loop?
redisresiliencedependencies - 57
After a kernel update, 18% of 2,000 RHEL hosts reboot repeatedly and five customer services lose quorum. How do you recover the fleet?
distributed-systemskernel - 58
The root filesystem on 80 Ubuntu API servers reaches 100% because one service writes 6 GB/minute of logs. What are your first actions?
storageapi - 59
An ext4 mail spool has 1.8 TB free but 0 inodes, so 12,000 messages cannot be queued. How do you restore service?
storagedata-structures - 60
PostgreSQL WAL fills a 2 TB volume at 180 GB/hour because a replication slot has been inactive for 11 hours. The database has 25 minutes before full. What do you do?
databasepostgresreplication - 61
A Java service deleted its 140 GB active log, but df still shows the filesystem at 98% while du reports only 40%. How do you reclaim the space?
storage - 62
A debug flag makes systemd-journald ingest 900 MB/minute across 300 hosts, and central logging is falling 45 minutes behind. How do you contain it?
logginglinuxsystem-design - 63
An XFS data volume is 99% full, LVM has 800 GB free, and writes will stop in about 40 minutes. How do you extend it without hiding the capacity failure?
capacitystoragecapacity-planning - 64
A backup repository reaches 97% during the weekly full, and 240 running jobs could consume the last 30 TB. What do you stop and what do you keep?
backupsbackup - 65
After a DNS change, 35% of clients receive SERVFAIL for an internal API while direct queries to both BIND authorities succeed. How do you localize it?
dnsnetwork-servicesapi - 66
Traffic to a 10.40.0.0/16 service fails from one data center after a route policy change, but outbound SYN packets leave normally. What do you check?
routing - 67
API requests show 2.8% packet loss between two VLANs, while intermediate MTR hops report values from 0% to 40%. How do you find the real loss point?
troubleshootingapi - 68
After enabling a VPN tunnel, small SSH commands work but 1,500-byte file transfers stall on 60% of paths. How do you handle the incident?
soft-skillsincidentsincident-management - 69
A Linux NAT gateway drops new connections at 120,000 sessions, conntrack is 99% full, and existing sessions still work. What do you do?
gatewaynetworkingnetwork-services - 70
Two servers answer for the same production IP, causing ARP entries to alternate every 20 seconds and 50% request failure. How do you recover?
protocols - 71
A four-link LACP bond remains up after one member degrades, but packet loss reaches 6% only for some flows. How do you isolate it?
troubleshooting - 72
Auditd shows a successful root SSH login from an unknown address 18 minutes ago on a production bastion. What do you do first?
sshremote-access - 73
An EDR alert finds a web shell on 14 of 600 Ubuntu web servers, and all 14 used the same image version. How do you respond?
alerting - 74
A Terraform service-account token was posted in a public repository for 47 minutes and can modify 12 production subscriptions. What is your sequence?
tokensterraform - 75
Forty Linux hosts jump to 100% CPU overnight, and an unknown process is mining cryptocurrency from /tmp. How do you investigate without losing scope?
linuxconcurrency - 76
File rename rates spike on 25 Windows servers, three backup shares become unreachable, and ransomware is suspected. How do you contain it?
backupsbackup - 77
A new sudo rule accidentally grants passwordless root to 1,200 developer accounts for 22 minutes. How do you respond?
passwordspermissionslinux - 78
An attacker changes 18 Active Directory DNS records and creates a Domain Admin account. Two domain controllers may be compromised. What is your recovery boundary?
dnsnetwork-services - 79
A 12-disk RAID 5 loses one 10 TB disk, rebuild is estimated at 31 hours, and read errors appear on a second disk. What do you do?
storageestimation - 80
A RAID 6 array is rebuilding after one disk failure when a second disk reports 38 pending sectors. The service still runs. What is your decision?
storage - 81
After a power event, an mdadm RAID 1 root array boots from one member and marks the other removed on 90 servers. How do you recover the fleet?
storage - 82
SAN latency rises from 3 ms to 180 ms, multipath changes paths 40 times per minute, and 70 virtual machines freeze. How do you stabilize storage?
virtualizationlatency - 83
A Ceph cluster reaches 91% raw use while recovering 180 degraded placement groups, and client latency has tripled. What do you do?
latency - 84
XFS reports metadata corruption on a 24 TB volume and remounts read-only during peak traffic. How do you recover it?
- 85
An NFS filer fails over in 45 seconds, but 600 clients remain stuck with stale file handles and jobs cannot resume. What is your recovery plan?
recoveryproblem-solving - 86
A RHEL 8 to RHEL 9 upgrade leaves 60 of 400 servers at an emergency shell because the root LVM volume is not found. How do you roll back?
storagerollback - 87
During a Windows Server DNS migration, query failures reach 28% after 2,000 clients receive the new resolver through DHCP. How do you reverse it?
dnsnetwork-servicesqueries - 88
A 300 TB storage migration reports 0.07% checksum mismatches after applications have written to the new array for 35 minutes. What do you do?
migrations - 89
An Ansible role migration changes SSH configuration on 800 servers, and 90 can no longer accept new sessions. How do you recover access?
configsshansible - 90
A VMware tools and virtual-hardware upgrade makes 40 of 250 VMs lose network connectivity after reboot. How do you roll back?
rollback - 91
A storage-controller firmware upgrade on the first node raises I/O p99 from 8 ms to 95 ms, but the vendor recommends completing all four nodes. What do you decide?
procurement - 92
CPU demand is growing 9% per month, 600 hosts are already at 72% peak, and procurement takes 90 days. What do you do now?
procurement - 93
A 1.2 PB storage platform grows 18 TB/week, sits at 78%, and a new project will add 160 TB in six weeks. How do you avoid a full cluster?
- 94
A 20-host VMware cluster reaches 88% RAM after sales adds 120 VMs, and losing one host would require ballooning 900 GB. What is your response?
- 95
Daily backups now take 11 hours in an 8-hour window, while the 10 Gb/s link averages only 48%. Where do you look and what do you change?
backupsbackup - 96
A data hall has 140 kW usable power, runs at 126 kW peak, and a requested 20-server expansion adds 18 kW. How do you handle the capacity crisis?
soft-skillscapacitycapacity-planning - 97
During an inode incident, a junior administrator deletes an unverified spool directory and 8,000 queued jobs disappear. How do you lead the recovery and coach them?
incidentsincident-managementdata-structures - 98
A mid-level administrator proposes disabling SELinux on 300 RHEL servers to meet a release deadline in four hours. How do you mentor them and unblock the release?
mentoringestimationhardening - 99
A new on-call administrator receives 43 alerts in 20 minutes and starts restarting hosts without a hypothesis. How do you intervene and build judgment?
on-callalertinghypothesis-testing - 100
A sysadmin submits a Bash cleanup script that will run as root on 2,500 hosts and delete files older than seven days. How do you review and mentor before rollout?
mentoringscripting