Skip to content

System Administrator interview questions

100 real questions with model answers and explanations for Senior candidates.

See a System Administrator resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

designlinux

I would make each hall an independent failure domain and reserve N+1 capacity outside it.

  • Spread every 3-node service across 3 halls, with separate power, top-of-rack switches, and storage paths per RHEL host.
  • Size 2 surviving halls for 100% load at 70% CPU, so normal placement cannot exceed 46% per hall.
  • Accept only designs whose dependency map and quarterly isolation test show 1 hall can disappear without shared DNS or identity loss.

Why interviewers ask this: The interviewer is checking whether redundancy covers correlated dependencies and includes enough surviving capacity.

availability

I would use interchangeable Ubuntu instances behind 2 independent HAProxy pairs and externalize all durable state.

  • Keep sessions in Redis and files in object storage, then require any 1 of 4 zones to lose 100% capacity safely.
  • At 250 requests/s per host and 60% target utilization, provision 800 hosts because 120,000/(250×0.60)=800.
  • Drain connections for 30 seconds and accept a release only when 5% canaries hold p99 below 200 ms with 25% headroom.

Why interviewers ask this: A strong answer combines statelessness, failure tolerance, capacity math, and measurable rollout gates.

designfailoverhigh-availability

I would run 1 active VRRP owner and 1 standby, with priority changes tied to HAProxy readiness.

  • Give node A priority 150 and node B 140, use unicast VRRP, and track HAProxy with a 20-point penalty.
  • Set advertisements to 500 ms, disable preemption after failback, and send 3 gratuitous ARP bursts when the VIP moves.
  • Keep 50% spare connection capacity per node and accept the design only if 20 failover trials restore traffic within 3 seconds.

Why interviewers ask this: The interviewer expects correct VRRP ownership behavior without unsafe failback or undersized nodes.

designlinux

I would require 3 of 5 Corosync votes and successful STONITH before Pacemaker starts the resource elsewhere.

  • Place 2 nodes in room A, 2 in B, and 1 in C, using 2 independent Corosync networks.
  • Configure IPMI fencing with 2 power paths and a 60-second timeout; quorum alone cannot stop an isolated writer.
  • Allow service only with at least 3 votes, and accept fencing after 20 tests prove the old node loses power before promotion.

Why interviewers ask this: The answer must distinguish majority authority from the fencing that prevents concurrent writers.

latency

I would choose active-passive ownership because 12 ms synchronous cross-site writes would directly tax all 8,000 writes/s.

  • Keep 1 primary site and 1 warm standby, replicate asynchronously, and cap observed lag at the 30-second RPO.
  • Fence the old primary through storage and power before promotion, accepting unused standby compute to avoid 2 writable authorities.
  • Size each site for 100% load at 65% CPU and require quarterly switchover to meet the 10-minute RTO.

Why interviewers ask this: The interviewer is evaluating consistency, latency, fencing, and the cost of passive capacity.

designdnsnetwork-services

I would separate authoritative zones from recursive resolution and deploy both across all 4 sites.

  • Run 2 Unbound resolvers per site with anycast or DHCP distribution, keeping 8 caches independent of authoritative BIND servers.
  • Use 30-second TTLs only for movable service records and 3,600 seconds for stable hosts, trading query load for convergence selectively.
  • Accept the design when 1 entire site can fail and 99.9% of uncached lookups still finish within 100 ms.

Why interviewers ask this: A senior design separates DNS roles and chooses TTLs from measured convergence and load requirements.

linux

I would use a 5-server Consul control plane and DNS names that hide ephemeral node addresses.

  • Spread 5 Consul servers across 3 failure domains, require 3 votes, and keep agents off the voting path.
  • Register readiness checks every 10 seconds with a 30-second deregistration threshold, accepting brief staleness instead of rapid endpoint flapping.
  • Cache results for 5 seconds and accept the design at 1,000 changes/minute with p99 lookup below 50 ms after 1 voter fails.

Why interviewers ask this: The interviewer is checking quorum placement, health semantics, cache convergence, and control-plane capacity.

designimmutability

I would rebuild versioned Packer images and replace hosts rather than patching 2,000 long-lived instances in place.

  • Pin the RHEL base digest and repository snapshot, sign each image, and retain 2 tested versions for rollback.
  • Promote through 1%, 10%, and 25% rings, requiring 30 minutes of health at each ring before fleet replacement.
  • Accept an image only after OpenSCAP, boot, systemd, and application tests pass, with 10% spare capacity for rolling replacement within 48 hours.

Why interviewers ask this: The answer should show reproducible artifacts, staged replacement, security evidence, and enough rollout capacity.

terraformpartitioning

I would create 36 remote state roots aligned to environment and lifecycle boundaries, not 1 global state.

  • Store all 36 states in a versioned encrypted backend with locking, and grant each CI identity access to only 1 state root.
  • Run refresh-only plans every 6 hours; reconcile drift through code unless a 1-hour break-glass record explicitly authorizes it.
  • Accept partitioning when a failed apply locks at most 1 domain and cross-state consumers use fewer than 10 stable outputs.

Why interviewers ask this: The interviewer is checking that state boundaries limit privilege and blast radius without creating uncontrolled coupling.

ansiblelinuxautomation

I would derive dynamic inventory from the CMDB and execute pinned Ansible environments through regional controllers.

  • Cache inventory for 5 minutes, group by region and service metadata, and abort if targeting differs by more than 2% from CMDB counts.
  • Use 5 regional controllers with 50 forks each, keeping WAN traffic local and limiting any play to 1 service group.
  • Accept a playbook only after check mode and idempotency tests show the 2nd run changes 0 resources on 100 representative hosts.

Why interviewers ask this: A strong answer makes inventory authoritative, execution bounded, and idempotency objectively testable.

ansibleautomationkernel

I would use serial push waves for the risky change and reserve pull mode for routine low-risk convergence.

  • Start with 18 hosts, then 180, then 20% waves, setting max_fail_percentage to 2 and pausing 15 minutes.
  • Use the Ansible sysctl module to update running and persistent state without a reboot, and require the 2nd execution to report 0 changes.
  • Stop expansion if health checks exceed 1% failures; push gives wave control, while 30-minute pull convergence reduces central-controller load later.

Why interviewers ask this: The interviewer wants a deliberate push-versus-pull choice, bounded waves, and real idempotency gates.

estimationdesign

I would promote signed repository snapshots through 4 rings so every host installs the same tested package set.

  • Mirror RHEL and Ubuntu upstreams daily, retain 3 immutable snapshots, and reject packages whose GPG signatures or SBOM checks fail.
  • Patch 1%, 10%, 30%, then 59% of hosts with 24-hour observation between rings, completing within 14 days.
  • Accept each ring below 0.5% service failures and keep 15% capacity spare, trading slower exposure reduction for controlled blast radius.

Why interviewers ask this: The answer should connect repository provenance, deterministic versions, rollout math, and explicit safety thresholds.

system-designdesignstorage

I would model hard requirements separately from startup order and avoid making 900 services depend on transient network reachability.

  • Add RequiresMountsFor=/srv/data plus Wants= and After=network-online.target, but let the application retry Vault for 120 seconds instead of ordering on it.
  • Use Restart=on-failure, RestartSec=5, and StartLimitBurst=6 per 60 seconds to prevent a fleet-wide restart storm.
  • Accept the unit after 50 boot tests prove data is mounted before start and shutdown completes within the 30-second TimeoutStopSec.

Why interviewers ask this: The interviewer is checking precise systemd dependency semantics and bounded behavior around external services.

designresilience

I would place at least 2 writable Windows Server domain controllers in each office and make sites topology-aware.

  • Run DNS and Global Catalog on all 6 controllers, map 3 AD Sites to subnets, and avoid a remote authentication dependency.
  • Keep all 5 FSMO roles assigned deliberately across 2 central controllers; roles are movable ownership, not active-active redundancy.
  • Accept the design when loss of 1 office leaves logon p95 below 2 seconds and replication converges within 15 minutes.

Why interviewers ask this: A senior answer distinguishes domain-controller redundancy, site locality, DNS, and FSMO role ownership.

monitoringlinux

I would design for 200,000 samples/s because 2,000×1,500/15 equals 200,000 before headroom.

  • Provision 2 identical Prometheus replicas per regional shard independently and size each for 300,000 samples/s, adding 50% ingestion headroom.
  • Keep 15 days locally, remote-write to Thanos, and alert when WAL growth or ingestion exceeds 70% of tested capacity.
  • Accept the platform after a 24-hour load test sustains 300,000 samples/s with query p99 below 2 seconds.

Why interviewers ask this: The interviewer is evaluating sample-rate math, HA duplication, retention design, and measurable headroom.

designavailabilitymonitoring

I would shard collection by region, duplicate each shard, and use Thanos for deduplicated 30-day global queries.

  • Deploy 2 Prometheus replicas in each of 5 regions with external labels, so 10 collectors survive 1 replica loss.
  • Send blocks every 2 hours to object storage and retain 24 hours locally, accepting remote-query latency for older data.
  • Accept regional isolation when 1 WAN link fails, local alerts still evaluate within 30 seconds, and global query p95 remains below 5 seconds.

Why interviewers ask this: The design must preserve local monitoring while providing durable, deduplicated global visibility.

monitoringalerting

I would impose label budgets before Grafana dashboards and route only actionable alerts through Alertmanager.

  • Limit each exporter to 2,000 active series and reject labels such as user_id that can create more than 10,000 values.
  • Route severity 1 to PagerDuty within 60 seconds and severity 2 to team queues, grouped by service for 5 minutes.
  • Accept a rule only with 1 owner, 1 runbook, and less than 5% false pages across a 30-day review.

Why interviewers ask this: The interviewer expects concrete controls for cardinality cost and alert actionability rather than more dashboards.

retentiondesignlogging

I would buffer logs at the edge, index selected fields centrally, and tier 240 TB of 30-day data.

  • Run Vector on 4,000 hosts with 10 GB disk buffers, protecting applications when OpenSearch is unavailable for 2 hours.
  • Keep 7 days on SSD-backed OpenSearch and 23 days in object storage, accepting slower searches for older records.
  • Cap indexed fields at 50 per source and accept the design when 16 TB/day replay keeps ingestion below 70% capacity.

Why interviewers ask this: A strong design balances buffering, searchable retention, indexing cost, and replay capacity.

I would encode CIS Level 1 as OpenSCAP policy and keep SELinux enforcing with 3 application-specific domains.

  • Apply 95% of the CIS profile automatically, documenting time-limited exceptions for the remaining 5% with 90-day expiry.
  • Generate policy from audited denials, map 3 ports to app_port_t, label /srv/app app_var_lib_t, and reject broad allow rules or permissive domains.
  • Accept a host at 98% scan compliance with 0 unexpected AVC denials across a 24-hour workload test.

Why interviewers ask this: The interviewer is checking measurable baseline enforcement and narrow SELinux policy rather than disabling controls.

hardeningdesignconcurrency

I would deploy 2 enforcing AppArmor profiles that grant each daemon only its required files, capabilities, and network family.

  • Allow daemon A /srv/a/** r, 1 writable log path, network inet stream, and capability net_bind_service; nftables exposes only TCP 443.
  • Run profiles in complain mode for 48 hours, review denials, then switch 10%, 50%, and 100% rings to enforce.
  • Accept enforcement with 0 unexpected denials and p99 latency growth below 2% at 20,000 requests/s.

Why interviewers ask this: A strong answer uses concrete AppArmor path rules and measured promotion rather than a generic security claim.

Locked questions

  • 21

    How would you design auditd collection for 3,500 Linux hosts while limiting audit overhead to 3% CPU?

    designlinux
  • 22

    Design SSH and PAM access for 1,800 Linux hosts where 60 administrators need privileged sessions within 5 minutes.

    sshlinuxremote-access
  • 23

    How would you provide secrets to 2,200 Linux services with 24-hour rotation and no static credentials in images?

    secretslinux
  • 24

    Design network segmentation for 5,000 servers in 4 trust zones while permitting 300 documented service flows.

    design
  • 25

    How would you design LVM for a 40 TB database host needing 2-hour maintenance snapshots and 20% growth headroom?

    designstoragedatabase
  • 26

    Choose RAID for 12 disks of 8 TB each when the workload needs 70 TB usable and any 2-disk failure tolerance.

    storage
  • 27

    Design a filesystem for 300 million 8 KB files requiring 25,000 random IOPS and 30 TB usable capacity.

    designcapacitystorage
  • 28

    How would you design NFS for 1,200 Linux clients needing 12 GB/s aggregate reads and 99.95% availability?

    designavailabilityaggregation
  • 29

    Design distributed storage for 5 PB raw capacity across 60 Linux nodes while tolerating loss of 1 entire rack.

    distributeddesigncapacity
  • 30

    How would you implement 3-2-1 backup for 800 servers containing 600 TB with a 24-hour RPO?

    backupsdisaster-recoverybackup
  • 31

    Size a backup path that must transfer 24 TB during an 8-hour window while retaining 25% throughput headroom.

    backupsbackupthroughput
  • 32

    Design backup validation for 200 critical servers with a 15-minute RPO and 2-hour RTO.

    designbackupsdisaster-recovery
  • 33

    How would you design warm-standby DR for 1,000 services across 2 regions with a 5-minute RPO and 30-minute RTO?

    designdisaster-recovery
  • 34

    Design quorum and replication for a 2-site service that needs 0 split brain and can add 1 witness location.

    designdistributed-systemsreplication
  • 35

    Plan CPU capacity for 2,400 hosts with 16 cores, 45% peak use, 30% growth, and loss of 1 of 3 zones.

    capacitycapacity-planning
  • 36

    Plan RAM for 1,000 virtualization guests averaging 12 GB with 20% growth and a 1-host maintenance reserve per 20-host cluster.

    virtualization
  • 37

    Size storage for 600 TB of data growing 4 TB/week for 26 weeks with 2 total copies and 20% free space.

  • 38

    Design a VMware cluster for 600 VMs averaging 4 vCPUs across 12 hosts with N+2 tolerance and a 4:1 overcommit limit.

    design
  • 39

    How would you design KVM virtualization for 300 VMs on 10 hosts with 25 Gb/s networking and 1-host maintenance tolerance?

    designvirtualization
  • 40

    Design Docker isolation for 200 containers on 20 Linux hosts with a 70% CPU utilization target.

    dockercontainerslinux
  • 41

    How would you secure 1,000 Docker workloads that expose 80 services and require 10 writable paths?

    docker
  • 42

    Design a small Kubernetes worker layer for 120 pods across 6 nodes with 1-node failure tolerance.

    designkubernetes
  • 43

    How would you provide Kubernetes storage for 40 stateful pods needing 2 TB each, a 15-minute RPO, and a 2-hour RTO?

    disaster-recoverykubernetes
  • 44

    Design Layer 4 load balancing for 500 TCP services carrying 200,000 concurrent connections across 3 zones.

    designload-balancingnetworking
  • 45

    How would you design routed management networks for 3,000 hosts across 6 sites with 2 independent WAN paths?

    designrouting
  • 46

    Design time synchronization for 5,000 Linux and Windows hosts requiring clocks within 50 ms across 5 sites.

    designlinux
  • 47

    How would you design centralized Linux identity for 2,700 hosts and 8,000 users with 1-site failure tolerance?

    designlinux
  • 48

    Design lifecycle management for 500 FreeBSD and 2,500 Linux hosts with 3 supported OS generations.

    designlinux
  • 49

    How would you manage configuration for 3,200 mostly Linux hosts and 400 intermittently connected Windows servers?

    configlinux
  • 50

    Design a 3,000-host platform combining HA, patching, monitoring, backup, and security with a 99.95% target.

    designbackupsbackup
  • 51

    A 32-core RHEL application host jumps from load 18 to 180, while CPU utilization is only 35% and 12 services time out. How do you diagnose and stabilize it?

  • 52

    A 128 GB Ubuntu database host has 96% RAM use, 42 GB of swap occupied, kswapd at 70% CPU, and query p99 rising from 80 ms to 4 seconds. What do you do?

    databasequeriesmemory
  • 53

    At 09:20, authentication starts failing on 700 of 900 servers, but CPU and memory on the application hosts remain normal. How do you find the shared failure?

    authmemory
  • 54

    A faulty unit starts 25,000 processes on a 64-core host, PID usage reaches 99%, and SSH can barely fork. How do you regain control?

    sshremote-accessconcurrency
  • 55

    Thirty web nodes show load 60 on 16 cores, 55% iowait, and a shared storage latency jump from 4 ms to 220 ms. What is your incident decision?

    incidentsincident-managementlatency
  • 56

    A Redis dependency recovers after 8 minutes, then 400 application instances retry at once and drive it back to 100% CPU. How do you break the recovery loop?

    redisresiliencedependencies
  • 57

    After a kernel update, 18% of 2,000 RHEL hosts reboot repeatedly and five customer services lose quorum. How do you recover the fleet?

    distributed-systemskernel
  • 58

    The root filesystem on 80 Ubuntu API servers reaches 100% because one service writes 6 GB/minute of logs. What are your first actions?

    storageapi
  • 59

    An ext4 mail spool has 1.8 TB free but 0 inodes, so 12,000 messages cannot be queued. How do you restore service?

    storagedata-structures
  • 60

    PostgreSQL WAL fills a 2 TB volume at 180 GB/hour because a replication slot has been inactive for 11 hours. The database has 25 minutes before full. What do you do?

    databasepostgresreplication
  • 61

    A Java service deleted its 140 GB active log, but df still shows the filesystem at 98% while du reports only 40%. How do you reclaim the space?

    storage
  • 62

    A debug flag makes systemd-journald ingest 900 MB/minute across 300 hosts, and central logging is falling 45 minutes behind. How do you contain it?

    logginglinuxsystem-design
  • 63

    An XFS data volume is 99% full, LVM has 800 GB free, and writes will stop in about 40 minutes. How do you extend it without hiding the capacity failure?

    capacitystoragecapacity-planning
  • 64

    A backup repository reaches 97% during the weekly full, and 240 running jobs could consume the last 30 TB. What do you stop and what do you keep?

    backupsbackup
  • 65

    After a DNS change, 35% of clients receive SERVFAIL for an internal API while direct queries to both BIND authorities succeed. How do you localize it?

    dnsnetwork-servicesapi
  • 66

    Traffic to a 10.40.0.0/16 service fails from one data center after a route policy change, but outbound SYN packets leave normally. What do you check?

    routing
  • 67

    API requests show 2.8% packet loss between two VLANs, while intermediate MTR hops report values from 0% to 40%. How do you find the real loss point?

    troubleshootingapi
  • 68

    After enabling a VPN tunnel, small SSH commands work but 1,500-byte file transfers stall on 60% of paths. How do you handle the incident?

    soft-skillsincidentsincident-management
  • 69

    A Linux NAT gateway drops new connections at 120,000 sessions, conntrack is 99% full, and existing sessions still work. What do you do?

    gatewaynetworkingnetwork-services
  • 70

    Two servers answer for the same production IP, causing ARP entries to alternate every 20 seconds and 50% request failure. How do you recover?

    protocols
  • 71

    A four-link LACP bond remains up after one member degrades, but packet loss reaches 6% only for some flows. How do you isolate it?

    troubleshooting
  • 72

    Auditd shows a successful root SSH login from an unknown address 18 minutes ago on a production bastion. What do you do first?

    sshremote-access
  • 73

    An EDR alert finds a web shell on 14 of 600 Ubuntu web servers, and all 14 used the same image version. How do you respond?

    alerting
  • 74

    A Terraform service-account token was posted in a public repository for 47 minutes and can modify 12 production subscriptions. What is your sequence?

    tokensterraform
  • 75

    Forty Linux hosts jump to 100% CPU overnight, and an unknown process is mining cryptocurrency from /tmp. How do you investigate without losing scope?

    linuxconcurrency
  • 76

    File rename rates spike on 25 Windows servers, three backup shares become unreachable, and ransomware is suspected. How do you contain it?

    backupsbackup
  • 77

    A new sudo rule accidentally grants passwordless root to 1,200 developer accounts for 22 minutes. How do you respond?

    passwordspermissionslinux
  • 78

    An attacker changes 18 Active Directory DNS records and creates a Domain Admin account. Two domain controllers may be compromised. What is your recovery boundary?

    dnsnetwork-services
  • 79

    A 12-disk RAID 5 loses one 10 TB disk, rebuild is estimated at 31 hours, and read errors appear on a second disk. What do you do?

    storageestimation
  • 80

    A RAID 6 array is rebuilding after one disk failure when a second disk reports 38 pending sectors. The service still runs. What is your decision?

    storage
  • 81

    After a power event, an mdadm RAID 1 root array boots from one member and marks the other removed on 90 servers. How do you recover the fleet?

    storage
  • 82

    SAN latency rises from 3 ms to 180 ms, multipath changes paths 40 times per minute, and 70 virtual machines freeze. How do you stabilize storage?

    virtualizationlatency
  • 83

    A Ceph cluster reaches 91% raw use while recovering 180 degraded placement groups, and client latency has tripled. What do you do?

    latency
  • 84

    XFS reports metadata corruption on a 24 TB volume and remounts read-only during peak traffic. How do you recover it?

  • 85

    An NFS filer fails over in 45 seconds, but 600 clients remain stuck with stale file handles and jobs cannot resume. What is your recovery plan?

    recoveryproblem-solving
  • 86

    A RHEL 8 to RHEL 9 upgrade leaves 60 of 400 servers at an emergency shell because the root LVM volume is not found. How do you roll back?

    storagerollback
  • 87

    During a Windows Server DNS migration, query failures reach 28% after 2,000 clients receive the new resolver through DHCP. How do you reverse it?

    dnsnetwork-servicesqueries
  • 88

    A 300 TB storage migration reports 0.07% checksum mismatches after applications have written to the new array for 35 minutes. What do you do?

    migrations
  • 89

    An Ansible role migration changes SSH configuration on 800 servers, and 90 can no longer accept new sessions. How do you recover access?

    configsshansible
  • 90

    A VMware tools and virtual-hardware upgrade makes 40 of 250 VMs lose network connectivity after reboot. How do you roll back?

    rollback
  • 91

    A storage-controller firmware upgrade on the first node raises I/O p99 from 8 ms to 95 ms, but the vendor recommends completing all four nodes. What do you decide?

    procurement
  • 92

    CPU demand is growing 9% per month, 600 hosts are already at 72% peak, and procurement takes 90 days. What do you do now?

    procurement
  • 93

    A 1.2 PB storage platform grows 18 TB/week, sits at 78%, and a new project will add 160 TB in six weeks. How do you avoid a full cluster?

  • 94

    A 20-host VMware cluster reaches 88% RAM after sales adds 120 VMs, and losing one host would require ballooning 900 GB. What is your response?

  • 95

    Daily backups now take 11 hours in an 8-hour window, while the 10 Gb/s link averages only 48%. Where do you look and what do you change?

    backupsbackup
  • 96

    A data hall has 140 kW usable power, runs at 126 kW peak, and a requested 20-server expansion adds 18 kW. How do you handle the capacity crisis?

    soft-skillscapacitycapacity-planning
  • 97

    During an inode incident, a junior administrator deletes an unverified spool directory and 8,000 queued jobs disappear. How do you lead the recovery and coach them?

    incidentsincident-managementdata-structures
  • 98

    A mid-level administrator proposes disabling SELinux on 300 RHEL servers to meet a release deadline in four hours. How do you mentor them and unblock the release?

    mentoringestimationhardening
  • 99

    A new on-call administrator receives 43 alerts in 20 minutes and starts restarting hosts without a hypothesis. How do you intervene and build judgment?

    on-callalertinghypothesis-testing
  • 100

    A sysadmin submits a Bash cleanup script that will run as root on 2,500 hosts and delete files older than seven days. How do you review and mentor before rollout?

    mentoringscripting