Network Engineer interview questions
100 real questions with model answers and explanations for Senior candidates.
See a Network Engineer resume example →Practice with flashcards
Spaced repetition · Hunter Pass
Questions
I would build 6 routed access blocks, each with 2 distribution switches, and connect them to a redundant Layer 3 core.
- Each block contains 8 access switches, dual-homed with 2 routed links, so a VLAN failure affects at most 3,000 users.
- OSPF summaries per building keep roughly 120 access prefixes from flooding across all 6 blocks, trading granular remote paths for stability.
- I would validate sub-3-second recovery by removing 1 distribution node while measuring loss, routing convergence, and remaining 40 Gbps uplink headroom.
Why interviewers ask this: The interviewer is testing whether topology, fault containment, route scale, and recovery criteria form one coherent campus design.
I would start with a 2-node collapsed core only if each node can sustain 90 Gbps after 1-node loss.
- Today, 24 access switches with 2 x 10 Gbps uplinks fit a redundant pair, while eliminating 2 core switches reduces ports and routing adjacencies.
- The design reserves 50 access attachments and 30% control-plane headroom, because the combined tier owns policy and transit state.
- I would trigger a separate Layer 3 core at 40 switches or 70% port use, accepting an earlier migration to preserve failure-domain boundaries.
Why interviewers ask this: A strong answer chooses the simpler topology only with explicit capacity, growth, and redesign thresholds.
The healthy leaf is 1.5:1 oversubscribed because 1,200 Gbps of server access shares 800 Gbps of fabric uplinks.
- Losing 1 uplink leaves 700 Gbps, changing the ratio to about 1.71:1 and capping any 1 flow at 100 Gbps.
- I would model storage and backup windows separately, requiring their p95 aggregate below 560 Gbps, or 80% of degraded uplink capacity.
- Validation uses 8-way ECMP traffic with mixed 10 Gbps and 80 Gbps flows, accepting under 1% loss and p99 latency below 150 microseconds.
Why interviewers ask this: The interviewer is evaluating correct capacity math and whether degraded-state workload behavior drives acceptance.
I would give every leaf 8 x 100 Gbps uplinks, producing 12.8 Tbps of theoretical bisection capacity between 2 groups of 16 leaves.
- Each spine receives 32 x 100 Gbps leaf links and needs at least 3.2 Tbps nonblocking forwarding in both directions.
- After 1 spine fails, bisection falls to 11.2 Tbps, so the 12 Tbps SLA requires a 9th spine or lower admitted demand.
- I would test 2-hop paths at 90% load, requiring p99 below 180 microseconds and under 0.1% queue drops.
Why interviewers ask this: A strong answer computes healthy and failed bisection capacity instead of quoting nominal port bandwidth.
I would make the 2 paths independent across device, power, conduit, carrier, facility, and regional control-plane domains.
- Each data center uses 2 edge routers in separate rooms, fed by 2 power systems and carrier entrances verified on physical-route records.
- Regional summaries ensure losing 1 site withdraws only its 40 service prefixes, while the paired site retains at least 60% of regional capacity.
- Quarterly validation removes 1 router, 1 circuit, and 1 route reflector separately, requiring recovery under 5 seconds without exceeding 70% surviving utilization.
Why interviewers ask this: The interviewer is checking whether nominal redundancy is traced through hidden shared dependencies and tested in degraded capacity.
I would use area 0 as the transit backbone and place 6 groups of 100 routers in areas.
- 2 ABRs per area summarize site blocks, reducing 4,000 specifics to about 60 aggregates outside their origin area, with less precise remote failover.
- Point-to-point links and passive user interfaces limit adjacency count, while BFD near 300 milliseconds x 3 targets failure detection below 1 second.
- I would validate 1 ABR loss with SPF and LSA telemetry, requiring forwarding recovery below 3 seconds and 0 routing loops.
Why interviewers ask this: A strong answer combines OSPF hierarchy, addressing discipline, timer realism, and measured convergence.
I would place 100 routers in each of 8 Level 1 areas and connect redundant borders through a compact Level 2 backbone.
- Level 1 routers receive a default toward Level 2, while regional /16 summaries keep roughly 10,000 access prefixes from entering every LSP database.
- Wide metrics, point-to-point adjacencies, and BFD at 250 milliseconds x 3 support fast detection without aggressive IS-IS hello churn.
- Failure tests remove 1 Level 1/2 router and require recovery below 2 seconds, 0 transient loops, and under 70% CPU.
Why interviewers ask this: The interviewer is testing whether IS-IS hierarchy is tied to flooding scale, summarization, and measurable recovery.
I would use eBGP with private ASNs per leaf and a shared or distinct spine ASN, creating 576 simple point-to-point sessions.
- Each session carries only loopback /32 and link reachability, keeping tenant routes out of the underlay and policy under 20 lines.
- Unnumbered links reduce address inventory, but /31 links make packet captures and per-link operations easier, so I would choose /31 for this team.
- BFD at 200 milliseconds x 3 plus 6-way ECMP must restore traffic within 1 second after any 1 uplink fails.
Why interviewers ask this: A strong answer quantifies eBGP session scale and keeps the physical fabric deliberately separate from overlay policy.
I would deploy 2 route reflectors in separate sites and give every router a session to both, yielding 800 client-to-reflector sessions.
- Reflectors receive the same 200,000 prefixes and use identical import policy, but they do not sit in the forwarding path.
- Add-Path advertises 2 useful exits for critical families, trading extra RIB memory for less path hiding and faster failover.
- I would cap each client at 250,000 prefixes and validate loss of 1 reflector with no withdrawals and under 3-second best-path stabilization.
Why interviewers ask this: The interviewer is checking session math, reflector placement, path visibility, and failure validation.
I would generate explicit import and export policy from 2 reviewed prefix sets, with default deny at every neighbor boundary.
- Local preference 200 selects the primary outbound carrier and 150 the secondary; communities separate customer, peer, transit, and blackhole actions.
- Export permits only the 300 owned prefixes, while a precommit count gate blocks output above 350 before an internal leak becomes transit.
- Import allows up to 1,300,000 routes, rejects bogons and RPKI-invalid origins, and must pass 20 policy fixtures before deployment.
Why interviewers ask this: A strong answer makes routing intent bounded, machine-testable, and different for inbound and outbound control.
I would combine sub-1-second detection with preinstalled alternate next hops, rather than reprogramming 50,000 prefixes.
- BFD at 300 milliseconds x 3 detects the failed path in about 900 milliseconds, with conservative timers on unstable access circuits.
- BGP PIC edge and core retain 2 eligible next hops, trading additional FIB state for failover that does not scale with prefix count.
- Route reflectors advertise 2 paths with Add-Path, and tests require under 2 seconds of loss with fewer than 3 route changes per prefix.
Why interviewers ask this: The interviewer is evaluating detection, path visibility, forwarding repair, and stability as separate convergence stages.
I would isolate the 2 sessions in a dedicated VRF and accept only the partner's approved set of 300 exact prefixes.
- A maximum of 350 prefixes, default-deny export, and AS-path filters bound route volume and stop accidental transit through our ASN.
- TCP-AO or MD5, GTSM with TTL 255, and control-plane policing at 100 packets per second protect the sessions themselves.
- Both edges use the same generated policy, and 12 negative tests must reject unauthorized prefixes, private AS leakage, and a default route.
Why interviewers ask this: A strong answer protects both BGP session establishment and the routing information accepted through it.
I would create 1 VRF per customer on each serving PE and distribute VPNv4 or VPNv6 routes between PEs with MP-BGP.
- Unique RDs distinguish overlapping /24 routes, while import and export RTs express 120 private VPNs and any approved shared-service topology.
- The outer label reaches the egress PE and the inner label selects its VRF, keeping 960 customer attachments out of P-router tables.
- 2 route reflectors and 120,000-route PE limits provide headroom, while tests require 0 cross-customer reachability across 50 sampled VRFs.
Why interviewers ask this: The interviewer is checking correct separation of route uniqueness, VPN membership, and label forwarding at scale.
I would run SR-MPLS over the existing IGP and steer the 3 services with policies computed from latency and affinity constraints.
- Node SIDs provide shortest-path instructions, while adjacency SIDs force selected links; a maximum 8-label stack stays within validated hardware depth.
- TI-LFA precomputes repair paths and targets under 50 milliseconds, trading extra label and FIB state for local recovery before global convergence.
- A controller rejects paths above 25 milliseconds one-way or 70% utilization, and 10 failure tests verify no loop or policy fallback.
Why interviewers ask this: A strong answer connects SR path encoding to hardware limits, fast repair, and explicit constraint validation.
I would keep 136 loopback /32s in the Layer 3 underlay and carry 40,000 endpoint routes in MP-BGP EVPN.
- The 8 spine forwarding tables hold no tenant MAC or IP routes, reducing data-plane coupling while preserving transport through all 8 ECMP paths.
- 2 route reflectors distribute EVPN routes to 128 VTEPs and advertise 2 paths with Add-Path, trading RIB memory for path visibility.
- Validation churns 1,000 Type 2 routes and fails 1 underlay link; prefix count must stay stable and forwarding recover within 1 second.
Why interviewers ask this: The interviewer is checking strict route ownership between physical reachability and tenant control-plane state.
I would use 4 EVPN route functions so endpoint, segment, multicast-membership, and routed-prefix state remain explicit.
- Type 2 advertises 40,000 MAC and IP bindings, enabling host reachability and mobility without data-plane learning across all VTEPs.
- Type 3 advertises membership for 200 VNIs and builds BUM replication lists, trading control-plane state for bounded flooding.
- Type 5 carries 5,000 IP prefixes, while Types 1 and 4 coordinate Ethernet segments, aliasing, and designated forwarders for multihomed attachments.
Why interviewers ask this: A strong answer maps route types to concrete overlay state rather than treating EVPN as one generic advertisement.
I would assign 1 all-active Ethernet Segment Identifier per server bond and run LACP to both EVPN leaves.
- Type 1 and Type 4 routes advertise ESI membership, while remote VTEPs install 2 aliasing next hops for load sharing.
- Designated-forwarder election permits only 1 leaf to send each BUM flow toward the segment, and split horizon prevents reflection loops.
- Tests remove 1 leaf and require recovery below 2 seconds, no duplicate BUM frames above 0.01%, and all 200 bonds still forwarding.
Why interviewers ask this: The interviewer is checking EVPN multihoming mechanics, loop prevention, and measurable failover behavior.
I would instantiate the same Layer 3 gateway IP and MAC on each hosting leaf.
- Source hosts route at their local leaf, avoiding a 2-hop detour to centralized gateways and keeping the common p99 target of 250 microseconds feasible.
- Symmetric IRB uses 1 L3 VNI per tenant VRF, reducing remote MAC dependencies at the cost of additional VNI and Type 5 state.
- Templates keep each of 1,000 gateway definitions consistent across participating leaves, and validation checks 0 duplicate IPs plus unchanged gateway identity during 20 endpoint moves.
Why interviewers ask this: A strong answer ties distributed gateway placement to latency, IRB state, and configuration consistency.
I would not keep ingress replication because 119 remote copies turn 20 Mbps into 2.38 Gbps on the source leaf.
- EVPN Type 3 still advertises VTEP membership, while PIM-based underlay multicast moves copy replication into the fabric for this large VNI.
- Multicast adds RP, group, and troubleshooting state, so VNIs with 20 or fewer VTEPs remain on simpler ingress replication.
- ARP suppression and smaller broadcast domains must reduce BUM below 5% of any 100 Gbps uplink before acceptance.
Why interviewers ask this: The interviewer is testing replication math and whether complexity is introduced only beyond a clear scale threshold.
I would set underlay MTU to 9,050 bytes because IPv4 VXLAN adds about 50 bytes to a 9,000-byte inner frame.
- The 9,216-byte port limit leaves 166 bytes of margin, but every interconnect, port channel, and firewall path must support the same value.
- Tenant MTU remains 9,000, while control probes test 9,050-byte outer packets with DF set so fragmentation cannot hide a mismatch.
- Acceptance requires 0 fragmentation and 0 giant drops across all 8 ECMP paths between the 2 data centers.
Why interviewers ask this: A strong answer performs the encapsulation math and validates every device in the actual path.
Locked questions
- 21
A campus serves 20,000 users across 120 VLANs and 4 trust zones; how would you design VLAN, VRF, and ACL boundaries without stretching Layer 2?
designswitchingnetwork-security - 22
A data center has 3,000 workloads and 60 approved application flows; how would you implement default-deny microsegmentation without 3,000 IP-based ACLs?
- 23
An application zone peaks at 40 Gbps and 8 million sessions; how would you place 2 firewalls so any 1 node can fail without asymmetric state loss?
sessionsnetwork-securitynetworking - 24
A company has 12,000 users and 80 private applications across 3 regions; how would you design Zero Trust access without granting broad /16 reachability?
zero-trustdesign - 25
An SDN domain controls 500 switches and 1 million forwarding entries; how would you design a 3-controller cluster that survives 1 node loss?
designswitching - 26
A Cisco ACI fabric hosts 8,000 endpoints in 4 tenants, 60 EPGs, and 120 allowed relationships; how would you model policy and APIC resilience?
endpoints - 27
You need a source of truth for 700 devices, 20,000 interfaces, and 40,000 prefixes; what data model and quality gates would you design?
modelingdesigntypes - 28
You must deploy 12 configuration templates to 600 mixed-vendor devices with no more than 20 devices exposed per wave; how would you structure Ansible?
deploymentconfigansible - 29
A Python service must update BGP policy on 400 routers through NETCONF and 8 YANG models; how would you make transactions safe?
routingtransactionspython - 30
A GitOps repository drives 500 devices from 12 templates; how would you design a pipeline that limits a bad merge to 25 devices?
ci-cdgitopsdesign - 31
Before changing ACLs on 300 devices with 50,000 routes, which Batfish checks would you require for 120 intended application paths?
routing - 32
A public zone serves 25 million DNS queries per day with 20x traffic peaks; how would you design authoritative DNS across 4 sites?
designdnsnetwork-services - 33
A DNS anycast service uses 6 PoPs at 50,000 queries per second each and must steer around a failed PoP within 30 seconds; how would you route it?
dnsroutingnetwork-services - 34
An L4 service must handle 2 million concurrent TCP sessions and 200,000 new connections per second across 16 nodes; how would you balance it?
networkingprotocolssessions - 35
An HTTPS API receives 150,000 requests per second with a p99 budget of 30 milliseconds; how would you design 12 L7 proxies across 3 zones?
httpsdesign - 36
A service uses 8 ECMP paths at 100 Gbps and carries 500 Gbps with several 80 Gbps elephant flows; how would you avoid hot links after 1 path fails?
- 37
An anycast API runs in 12 PoPs with 2 Tbps aggregate demand and stateful logins lasting 30 minutes; how would you balance service traffic?
aggregationapi - 38
You must allocate addresses for 300 cloud VPCs or VNets and 40 on-premises sites with 30% growth reserve; how would you design hybrid IPAM?
design - 39
An AWS estate has 80 VPCs across 3 trust zones and 2 regions, with 10 Gbps east-west demand; how would you structure Transit Gateway routing?
gateway - 40
A hybrid design needs AWS Direct Connect and Azure ExpressRoute, each carrying 8 Gbps normally and 6 Gbps after 1 circuit fails; how would you build it?
design - 41
AWS and Azure exchange 30 Gbps of application traffic, but a central hub adds 70 milliseconds against a p99 target of 45 milliseconds; how would you design multi-cloud connectivity?
design - 42
A 250-branch network has 500 Mbps DIA and 100 Mbps LTE per branch; how would you design SD-WAN for voice under 150 milliseconds and 1% loss?
designwan - 43
You need encrypted spoke-to-spoke routing for 600 sites through 2 hub regions; how would you design DMVPN without 179,700 permanent tunnels?
encryptiondesign - 44
A hybrid WAN falls from 1 Gbps to 500 Mbps after failure; how would you allocate 4 QoS classes so voice and business traffic remain usable?
qos - 45
You need streaming telemetry from 800 devices, each exporting 60 metrics every 30 seconds; how would you size collection and retention?
retentionstreamingmonitoring - 46
A 4 x 100 Gbps backbone runs at p95 of 280 Gbps and grows 35% yearly; what upgrade preserves 30% headroom after 1 link fails?
- 47
An enterprise receives an IPv6 /32 for 200 sites and expects 20% growth; how would you allocate prefixes while keeping routing summaries stable?
ip-addressing - 48
A video platform has 40 multicast channels at 8 Mbps and 400 receivers across 20 sites; how would you avoid 128 Gbps of replicated unicast?
replication - 49
Your internet edge has 2 x 100 Gbps circuits, 80 Gbps normal demand, and a 2 Tbps attack design case; what DDoS architecture would you build?
design - 50
You must provide out-of-band access to 700 network devices across 12 sites during a total production-network loss; how would you design it?
design - 51
Traffic between VLAN 120 and VLAN 340 shows 2.5% ping loss and 18 ms jitter, while checkout errors rose to 4%; how do you localize the fault?
switching - 52
A 100 Gbps uplink averages 42% utilization, but storage goodput falls from 38 to 11 Gbps while 200-microsecond bursts recur every minute; what do you investigate?
- 53
After enabling VXLAN, 1,500-byte TCP sessions stall while 1,400-byte probes pass across 6 fabric paths; how do you diagnose the MTU failure?
networkingprotocolsdatacenter - 54
A 4-link 100 Gbps port channel carries 210 Gbps, yet one replication job drops from 75 to 24 Gbps after a member fails; how do you respond?
replication - 55
An eBGP session resets 14 times in 20 minutes and withdraws 180,000 routes, while the physical interface never drops; what is your incident plan?
incidentstypessessions - 56
A policy deployment leaks 26,000 internal routes to 2 transit providers and external monitors see them within 90 seconds; what do you do first?
deploymentmonitoringrouting - 57
A partner accidentally originates your /22 from AS 64580, attracting 35% of traffic despite your valid /22 announcement; how do you recover safely?
- 58
Users in 2 of 7 regions receive SERVFAIL for api.example.com after a delegation change, while direct queries to 3 authoritative servers work; how do you isolate it?
delegationapiqueries - 59
A DNSSEC key rollover causes 38% validation failures, and the zone's DNSKEY TTL is 900 seconds; what recovery sequence do you choose?
validation - 60
One of 8 DNS anycast PoPs returns stale records to 12% of clients, but its BGP announcement and health check remain green; how do you contain it?
dnsroutingnetwork-services - 61
Checkout RTT rises from 28 to 96 ms for 1 region after 02:10, but interface counters show 0 errors; how do you prove where the path changed?
types - 62
TCP uploads lose 3% of packets in one direction across a 4-hop WAN, while downloads are clean and MTR disagrees by hop; how do you localize it?
networkingprotocolsconflict - 63
After an ACL change, 14 hosts in VLAN 210 can reach a payroll subnet that should allow only 3 jump servers; how do you stop and explain the leak?
networkingip-addressingswitching - 64
Tenant A learns 620 Tenant B prefixes after an EVPN policy edit, although the VRFs use different VNIs; how do you contain the route leak?
routing - 65
An uplink fails, but HSRP leaves the VIP on the isolated active switch and 4,000 clients blackhole for 6 minutes; what do you fix?
switchinghigh-availability - 66
A VRRP pair fails over in 1 second, but 30% of sessions reset because return traffic still uses the old firewall path; how do you recover?
sessionsnetwork-securitynetworking - 67
OSPF neighbors on 12 links oscillate between EXSTART and FULL after a software upgrade, causing 9-second outages; how do you diagnose it?
routing - 68
A misconfigured redistribution creates 18,000 OSPF external LSAs and drives 40 routers above 85% CPU; what is your recovery order?
routingredis - 69
An IS-IS Level 2 router reports overload and 6 sites take a 45 ms detour after a memory alarm; how do you handle the event?
soft-skillsmemoryrouting - 70
A new trunk creates an STP loop, broadcast traffic reaches 18 Gbps, and 22 access switches hit 95% CPU; how do you restore the campus?
switching - 71
A 4-member LACP bundle shows all links up, but 25% of flows fail after one switch reloads; what specific fault do you suspect and test?
switching - 72
A 100 Gbps optic accumulates 900 CRC errors per minute only above 70% load, and FEC corrected errors rise 12x; how do you decide what to replace?
- 73
One EVPN tenant loses 240 endpoints after a template change, while 19 other tenants remain healthy; how do you bound and recover it?
endpoints - 74
A virtual machine triggers 46 EVPN MAC mobility events in 60 seconds between 2 leaves, causing intermittent 8-second outages; what do you do?
- 75
BUM traffic in VNI 50120 jumps from 200 Mbps to 14 Gbps and saturates 3 VTEPs after an endpoint rollout; how do you contain it?
endpoints - 76
Jumbo storage traffic fails on 2 of 8 VXLAN ECMP paths after a spine replacement, while 1,500-byte traffic is clean; what is your recovery plan?
recoverydatacenter - 77
After adding a 3rd internet exit, 18% of TCP sessions hit firewall out-of-state drops although both directions are reachable; how do you restore service?
networkingprotocolssessions - 78
A NAT gateway reaches 98% of its 64,000 ports for 1 public IP, and API connection failures climb to 17%; what do you change first?
gatewaynetworkingnetwork-services - 79
An L4 load balancer loses 2 of 12 nodes and resets 9% of 1.6 million TCP sessions despite healthy backend servers; how do you limit damage?
load-balancingnetworkingprotocols - 80
An L7 proxy deployment raises 5xx responses from 0.2% to 11% because health checks mark 40% of backends down; how do you decide rollback?
deploymenthealth-checksrollback - 81
A 620 Gbps UDP flood saturates 2 x 100 Gbps internet links within 3 minutes; what actions do you take before local ACLs can help?
networkingprotocols - 82
A 4 Mpps SYN flood drives an edge router to 92% CPU while links remain below 20 Gbps; how do you protect the control plane?
routing - 83
A campus DHCP scope has 0 free addresses for 6,200 clients, while the standby server reports PARTNER-DOWN for 18 minutes; how do you recover leases safely?
network-services - 84
After a switch upgrade, 1,800 IPv6 clients install a rogue default router with 30-second lifetime and lose SaaS access every 5 minutes; how do you stop it?
ip-addressingroutingswitching - 85
A 10 Gbps Direct Connect drops at 14:05, but VPN fallback delivers only 2 Gbps and 3 critical applications need 1.4 Gbps; how do you fail over?
network-security - 86
A Transit Gateway route-table edit sends 27 VPCs to a missing inspection attachment, raising timeout errors to 23%; how do you restore connectivity?
gatewayresiliencerouting - 87
An SD-WAN policy oscillates 36 branches between MPLS and DIA every 90 seconds when loss hovers near 1%; how do you stabilize service?
wan - 88
An MPLS VPN loses 8 sites because an LDP label is missing although the IGP route remains present; how do you isolate the forwarding break?
routingwannetwork-security - 89
On a 500 Mbps failed WAN path, voice uses 95 Mbps although its LLQ is configured for 15%, starving a 150 Mbps business class; what do you change?
config - 90
Automation pushes a wrong prefix list to 160 routers in 4 minutes, withdrawing 12,000 customer routes; how do you stop the rollout?
routing - 91
A configuration job succeeds on 73 of 100 switches, then loses API access and leaves mixed VLAN state; do you roll forward or back?
switchingapiconfig - 92
During a 4,000-user HSRP blackhole, a junior engineer wants to clear ARP on all access switches; how do you mentor them while restoring service?
protocolsswitchinghigh-availability - 93
A mid-level engineer sees 7 BGP resets in 10 minutes and proposes global dampening; how do you coach the investigation and containment?
routing - 94
A junior engineer replaces 2 optics and a fiber after seeing 300 CRC errors per minute, but the fault remains; how do you redirect the next 45 minutes?
react - 95
A 47-minute packet-loss incident ended after an ECMP member was drained, but no device alarm fired despite 6% checkout failures; what belongs in the postmortem?
incidents - 96
A DNS incident is mitigated in 8 minutes, but 9% of clients still fail for 1 hour because resolvers cached NXDOMAIN; how do you judge recovery?
dnsnetwork-servicescaching - 97
A route-policy change cuts latency from 70 to 35 ms but causes 2% loss on 1 carrier; do you roll back after 12 minutes or tune forward?
routinglatencyrollback - 98
IPsec tunnels at 60 branches drop every 55 minutes after a key-policy change, interrupting voice for 20 seconds; how do you isolate rekey failure?
network-security - 99
A telemetry change sends 180,000 packets per second to router CPUs, causing 24 OSPF adjacencies to flap; how do you recover the control plane?
routing - 100
At 09:00, 11 sites lose all applications and 3 recent network changes are plausible; how do you run the first 30 minutes and choose recovery?