Skip to content

Network Engineer interview questions

100 real questions with model answers and explanations for Senior candidates.

See a Network Engineer resume example

Practice with flashcards

Spaced repetition · Hunter Pass

Questions

designswitching

I would build 6 routed access blocks, each with 2 distribution switches, and connect them to a redundant Layer 3 core.

  • Each block contains 8 access switches, dual-homed with 2 routed links, so a VLAN failure affects at most 3,000 users.
  • OSPF summaries per building keep roughly 120 access prefixes from flooding across all 6 blocks, trading granular remote paths for stability.
  • I would validate sub-3-second recovery by removing 1 distribution node while measuring loss, routing convergence, and remaining 40 Gbps uplink headroom.

Why interviewers ask this: The interviewer is testing whether topology, fault containment, route scale, and recovery criteria form one coherent campus design.

switching

I would start with a 2-node collapsed core only if each node can sustain 90 Gbps after 1-node loss.

  • Today, 24 access switches with 2 x 10 Gbps uplinks fit a redundant pair, while eliminating 2 core switches reduces ports and routing adjacencies.
  • The design reserves 50 access attachments and 30% control-plane headroom, because the combined tier owns policy and transit state.
  • I would trigger a separate Layer 3 core at 40 switches or 70% port use, accepting an earlier migration to preserve failure-domain boundaries.

Why interviewers ask this: A strong answer chooses the simpler topology only with explicit capacity, growth, and redesign thresholds.

capacity

The healthy leaf is 1.5:1 oversubscribed because 1,200 Gbps of server access shares 800 Gbps of fabric uplinks.

  • Losing 1 uplink leaves 700 Gbps, changing the ratio to about 1.71:1 and capping any 1 flow at 100 Gbps.
  • I would model storage and backup windows separately, requiring their p95 aggregate below 560 Gbps, or 80% of degraded uplink capacity.
  • Validation uses 8-way ECMP traffic with mixed 10 Gbps and 80 Gbps flows, accepting under 1% loss and p99 latency below 150 microseconds.

Why interviewers ask this: The interviewer is evaluating correct capacity math and whether degraded-state workload behavior drives acceptance.

latency

I would give every leaf 8 x 100 Gbps uplinks, producing 12.8 Tbps of theoretical bisection capacity between 2 groups of 16 leaves.

  • Each spine receives 32 x 100 Gbps leaf links and needs at least 3.2 Tbps nonblocking forwarding in both directions.
  • After 1 spine fails, bisection falls to 11.2 Tbps, so the 12 Tbps SLA requires a 9th spine or lower admitted demand.
  • I would test 2-hop paths at 90% load, requiring p99 below 180 microseconds and under 0.1% queue drops.

Why interviewers ask this: A strong answer computes healthy and failed bisection capacity instead of quoting nominal port bandwidth.

design

I would make the 2 paths independent across device, power, conduit, carrier, facility, and regional control-plane domains.

  • Each data center uses 2 edge routers in separate rooms, fed by 2 power systems and carrier entrances verified on physical-route records.
  • Regional summaries ensure losing 1 site withdraws only its 40 service prefixes, while the paired site retains at least 60% of regional capacity.
  • Quarterly validation removes 1 router, 1 circuit, and 1 route reflector separately, requiring recovery under 5 seconds without exceeding 70% surviving utilization.

Why interviewers ask this: The interviewer is checking whether nominal redundancy is traced through hidden shared dependencies and tested in degraded capacity.

routing

I would use area 0 as the transit backbone and place 6 groups of 100 routers in areas.

  • 2 ABRs per area summarize site blocks, reducing 4,000 specifics to about 60 aggregates outside their origin area, with less precise remote failover.
  • Point-to-point links and passive user interfaces limit adjacency count, while BFD near 300 milliseconds x 3 targets failure detection below 1 second.
  • I would validate 1 ABR loss with SPF and LSA telemetry, requiring forwarding recovery below 3 seconds and 0 routing loops.

Why interviewers ask this: A strong answer combines OSPF hierarchy, addressing discipline, timer realism, and measured convergence.

designrouting

I would place 100 routers in each of 8 Level 1 areas and connect redundant borders through a compact Level 2 backbone.

  • Level 1 routers receive a default toward Level 2, while regional /16 summaries keep roughly 10,000 access prefixes from entering every LSP database.
  • Wide metrics, point-to-point adjacencies, and BFD at 250 milliseconds x 3 support fast detection without aggressive IS-IS hello churn.
  • Failure tests remove 1 Level 1/2 router and require recovery below 2 seconds, 0 transient loops, and under 70% CPU.

Why interviewers ask this: The interviewer is testing whether IS-IS hierarchy is tied to flooding scale, summarization, and measurable recovery.

I would use eBGP with private ASNs per leaf and a shared or distinct spine ASN, creating 576 simple point-to-point sessions.

  • Each session carries only loopback /32 and link reachability, keeping tenant routes out of the underlay and policy under 20 lines.
  • Unnumbered links reduce address inventory, but /31 links make packet captures and per-link operations easier, so I would choose /31 for this team.
  • BFD at 200 milliseconds x 3 plus 6-way ECMP must restore traffic within 1 second after any 1 uplink fails.

Why interviewers ask this: A strong answer quantifies eBGP session scale and keeps the physical fabric deliberately separate from overlay policy.

sessionsrouting

I would deploy 2 route reflectors in separate sites and give every router a session to both, yielding 800 client-to-reflector sessions.

  • Reflectors receive the same 200,000 prefixes and use identical import policy, but they do not sit in the forwarding path.
  • Add-Path advertises 2 useful exits for critical families, trading extra RIB memory for less path hiding and faster failover.
  • I would cap each client at 250,000 prefixes and validate loss of 1 reflector with no withdrawals and under 3-second best-path stabilization.

Why interviewers ask this: The interviewer is checking session math, reflector placement, path visibility, and failure validation.

ip-addressingrouting

I would generate explicit import and export policy from 2 reviewed prefix sets, with default deny at every neighbor boundary.

  • Local preference 200 selects the primary outbound carrier and 150 the secondary; communities separate customer, peer, transit, and blackhole actions.
  • Export permits only the 300 owned prefixes, while a precommit count gate blocks output above 350 before an internal leak becomes transit.
  • Import allows up to 1,300,000 routes, rejects bogons and RPKI-invalid origins, and must pass 20 policy fixtures before deployment.

Why interviewers ask this: A strong answer makes routing intent bounded, machine-testable, and different for inbound and outbound control.

designroutingnetwork-security

I would combine sub-1-second detection with preinstalled alternate next hops, rather than reprogramming 50,000 prefixes.

  • BFD at 300 milliseconds x 3 detects the failed path in about 900 milliseconds, with conservative timers on unstable access circuits.
  • BGP PIC edge and core retain 2 eligible next hops, trading additional FIB state for failover that does not scale with prefix count.
  • Route reflectors advertise 2 paths with Add-Path, and tests require under 2 seconds of loss with fewer than 3 route changes per prefix.

Why interviewers ask this: The interviewer is evaluating detection, path visibility, forwarding repair, and stability as separate convergence stages.

designrouting

I would isolate the 2 sessions in a dedicated VRF and accept only the partner's approved set of 300 exact prefixes.

  • A maximum of 350 prefixes, default-deny export, and AS-path filters bound route volume and stop accidental transit through our ASN.
  • TCP-AO or MD5, GTSM with TTL 255, and control-plane policing at 100 packets per second protect the sessions themselves.
  • Both edges use the same generated policy, and 12 negative tests must reject unauthorized prefixes, private AS leakage, and a default route.

Why interviewers ask this: A strong answer protects both BGP session establishment and the routing information accepted through it.

designroutingwan

I would create 1 VRF per customer on each serving PE and distribute VPNv4 or VPNv6 routes between PEs with MP-BGP.

  • Unique RDs distinguish overlapping /24 routes, while import and export RTs express 120 private VPNs and any approved shared-service topology.
  • The outer label reaches the egress PE and the inner label selects its VRF, keeping 960 customer attachments out of P-router tables.
  • 2 route reflectors and 120,000-route PE limits provide headroom, while tests require 0 cross-customer reachability across 50 sampled VRFs.

Why interviewers ask this: The interviewer is checking correct separation of route uniqueness, VPN membership, and label forwarding at scale.

routingwanlatency

I would run SR-MPLS over the existing IGP and steer the 3 services with policies computed from latency and affinity constraints.

  • Node SIDs provide shortest-path instructions, while adjacency SIDs force selected links; a maximum 8-label stack stays within validated hardware depth.
  • TI-LFA precomputes repair paths and targets under 50 milliseconds, trading extra label and FIB state for local recovery before global convergence.
  • A controller rejects paths above 25 milliseconds one-way or 70% utilization, and 10 failure tests verify no loop or policy fallback.

Why interviewers ask this: A strong answer connects SR path encoding to hardware limits, fast repair, and explicit constraint validation.

routingdatacenterendpoints

I would keep 136 loopback /32s in the Layer 3 underlay and carry 40,000 endpoint routes in MP-BGP EVPN.

  • The 8 spine forwarding tables hold no tenant MAC or IP routes, reducing data-plane coupling while preserving transport through all 8 ECMP paths.
  • 2 route reflectors distribute EVPN routes to 128 VTEPs and advertise 2 paths with Add-Path, trading RIB memory for path visibility.
  • Validation churns 1,000 Type 2 routes and fails 1 underlay link; prefix count must stay stable and forwarding recover within 1 second.

Why interviewers ask this: The interviewer is checking strict route ownership between physical reachability and tenant control-plane state.

routingdatacenterendpoints

I would use 4 EVPN route functions so endpoint, segment, multicast-membership, and routed-prefix state remain explicit.

  • Type 2 advertises 40,000 MAC and IP bindings, enabling host reachability and mobility without data-plane learning across all VTEPs.
  • Type 3 advertises membership for 200 VNIs and builds BUM replication lists, trading control-plane state for bounded flooding.
  • Type 5 carries 5,000 IP prefixes, while Types 1 and 4 coordinate Ethernet segments, aliasing, and designated forwarders for multihomed attachments.

Why interviewers ask this: A strong answer maps route types to concrete overlay state rather than treating EVPN as one generic advertisement.

design

I would assign 1 all-active Ethernet Segment Identifier per server bond and run LACP to both EVPN leaves.

  • Type 1 and Type 4 routes advertise ESI membership, while remote VTEPs install 2 aliasing next hops for load sharing.
  • Designated-forwarder election permits only 1 leaf to send each BUM flow toward the segment, and split horizon prevents reflection loops.
  • Tests remove 1 leaf and require recovery below 2 seconds, no duplicate BUM frames above 0.01%, and all 200 bonds still forwarding.

Why interviewers ask this: The interviewer is checking EVPN multihoming mechanics, loop prevention, and measurable failover behavior.

gatewaynetworkingip-addressing

I would instantiate the same Layer 3 gateway IP and MAC on each hosting leaf.

  • Source hosts route at their local leaf, avoiding a 2-hop detour to centralized gateways and keeping the common p99 target of 250 microseconds feasible.
  • Symmetric IRB uses 1 L3 VNI per tenant VRF, reducing remote MAC dependencies at the cost of additional VNI and Type 5 state.
  • Templates keep each of 1,000 gateway definitions consistent across participating leaves, and validation checks 0 duplicate IPs plus unchanged gateway identity during 20 endpoint moves.

Why interviewers ask this: A strong answer ties distributed gateway placement to latency, IRB state, and configuration consistency.

replicationkubernetes

I would not keep ingress replication because 119 remote copies turn 20 Mbps into 2.38 Gbps on the source leaf.

  • EVPN Type 3 still advertises VTEP membership, while PIM-based underlay multicast moves copy replication into the fabric for this large VNI.
  • Multicast adds RP, group, and troubleshooting state, so VNIs with 20 or fewer VTEPs remain on simpler ingress replication.
  • ARP suppression and smaller broadcast domains must reduce BUM below 5% of any 100 Gbps uplink before acceptance.

Why interviewers ask this: The interviewer is testing replication math and whether complexity is introduced only beyond a clear scale threshold.

datacenterperformance

I would set underlay MTU to 9,050 bytes because IPv4 VXLAN adds about 50 bytes to a 9,000-byte inner frame.

  • The 9,216-byte port limit leaves 166 bytes of margin, but every interconnect, port channel, and firewall path must support the same value.
  • Tenant MTU remains 9,000, while control probes test 9,050-byte outer packets with DF set so fragmentation cannot hide a mismatch.
  • Acceptance requires 0 fragmentation and 0 giant drops across all 8 ECMP paths between the 2 data centers.

Why interviewers ask this: A strong answer performs the encapsulation math and validates every device in the actual path.

Locked questions

  • 21

    A campus serves 20,000 users across 120 VLANs and 4 trust zones; how would you design VLAN, VRF, and ACL boundaries without stretching Layer 2?

    designswitchingnetwork-security
  • 22

    A data center has 3,000 workloads and 60 approved application flows; how would you implement default-deny microsegmentation without 3,000 IP-based ACLs?

  • 23

    An application zone peaks at 40 Gbps and 8 million sessions; how would you place 2 firewalls so any 1 node can fail without asymmetric state loss?

    sessionsnetwork-securitynetworking
  • 24

    A company has 12,000 users and 80 private applications across 3 regions; how would you design Zero Trust access without granting broad /16 reachability?

    zero-trustdesign
  • 25

    An SDN domain controls 500 switches and 1 million forwarding entries; how would you design a 3-controller cluster that survives 1 node loss?

    designswitching
  • 26

    A Cisco ACI fabric hosts 8,000 endpoints in 4 tenants, 60 EPGs, and 120 allowed relationships; how would you model policy and APIC resilience?

    endpoints
  • 27

    You need a source of truth for 700 devices, 20,000 interfaces, and 40,000 prefixes; what data model and quality gates would you design?

    modelingdesigntypes
  • 28

    You must deploy 12 configuration templates to 600 mixed-vendor devices with no more than 20 devices exposed per wave; how would you structure Ansible?

    deploymentconfigansible
  • 29

    A Python service must update BGP policy on 400 routers through NETCONF and 8 YANG models; how would you make transactions safe?

    routingtransactionspython
  • 30

    A GitOps repository drives 500 devices from 12 templates; how would you design a pipeline that limits a bad merge to 25 devices?

    ci-cdgitopsdesign
  • 31

    Before changing ACLs on 300 devices with 50,000 routes, which Batfish checks would you require for 120 intended application paths?

    routing
  • 32

    A public zone serves 25 million DNS queries per day with 20x traffic peaks; how would you design authoritative DNS across 4 sites?

    designdnsnetwork-services
  • 33

    A DNS anycast service uses 6 PoPs at 50,000 queries per second each and must steer around a failed PoP within 30 seconds; how would you route it?

    dnsroutingnetwork-services
  • 34

    An L4 service must handle 2 million concurrent TCP sessions and 200,000 new connections per second across 16 nodes; how would you balance it?

    networkingprotocolssessions
  • 35

    An HTTPS API receives 150,000 requests per second with a p99 budget of 30 milliseconds; how would you design 12 L7 proxies across 3 zones?

    httpsdesign
  • 36

    A service uses 8 ECMP paths at 100 Gbps and carries 500 Gbps with several 80 Gbps elephant flows; how would you avoid hot links after 1 path fails?

  • 37

    An anycast API runs in 12 PoPs with 2 Tbps aggregate demand and stateful logins lasting 30 minutes; how would you balance service traffic?

    aggregationapi
  • 38

    You must allocate addresses for 300 cloud VPCs or VNets and 40 on-premises sites with 30% growth reserve; how would you design hybrid IPAM?

    design
  • 39

    An AWS estate has 80 VPCs across 3 trust zones and 2 regions, with 10 Gbps east-west demand; how would you structure Transit Gateway routing?

    gateway
  • 40

    A hybrid design needs AWS Direct Connect and Azure ExpressRoute, each carrying 8 Gbps normally and 6 Gbps after 1 circuit fails; how would you build it?

    design
  • 41

    AWS and Azure exchange 30 Gbps of application traffic, but a central hub adds 70 milliseconds against a p99 target of 45 milliseconds; how would you design multi-cloud connectivity?

    design
  • 42

    A 250-branch network has 500 Mbps DIA and 100 Mbps LTE per branch; how would you design SD-WAN for voice under 150 milliseconds and 1% loss?

    designwan
  • 43

    You need encrypted spoke-to-spoke routing for 600 sites through 2 hub regions; how would you design DMVPN without 179,700 permanent tunnels?

    encryptiondesign
  • 44

    A hybrid WAN falls from 1 Gbps to 500 Mbps after failure; how would you allocate 4 QoS classes so voice and business traffic remain usable?

    qos
  • 45

    You need streaming telemetry from 800 devices, each exporting 60 metrics every 30 seconds; how would you size collection and retention?

    retentionstreamingmonitoring
  • 46

    A 4 x 100 Gbps backbone runs at p95 of 280 Gbps and grows 35% yearly; what upgrade preserves 30% headroom after 1 link fails?

  • 47

    An enterprise receives an IPv6 /32 for 200 sites and expects 20% growth; how would you allocate prefixes while keeping routing summaries stable?

    ip-addressing
  • 48

    A video platform has 40 multicast channels at 8 Mbps and 400 receivers across 20 sites; how would you avoid 128 Gbps of replicated unicast?

    replication
  • 49

    Your internet edge has 2 x 100 Gbps circuits, 80 Gbps normal demand, and a 2 Tbps attack design case; what DDoS architecture would you build?

    design
  • 50

    You must provide out-of-band access to 700 network devices across 12 sites during a total production-network loss; how would you design it?

    design
  • 51

    Traffic between VLAN 120 and VLAN 340 shows 2.5% ping loss and 18 ms jitter, while checkout errors rose to 4%; how do you localize the fault?

    switching
  • 52

    A 100 Gbps uplink averages 42% utilization, but storage goodput falls from 38 to 11 Gbps while 200-microsecond bursts recur every minute; what do you investigate?

  • 53

    After enabling VXLAN, 1,500-byte TCP sessions stall while 1,400-byte probes pass across 6 fabric paths; how do you diagnose the MTU failure?

    networkingprotocolsdatacenter
  • 54

    A 4-link 100 Gbps port channel carries 210 Gbps, yet one replication job drops from 75 to 24 Gbps after a member fails; how do you respond?

    replication
  • 55

    An eBGP session resets 14 times in 20 minutes and withdraws 180,000 routes, while the physical interface never drops; what is your incident plan?

    incidentstypessessions
  • 56

    A policy deployment leaks 26,000 internal routes to 2 transit providers and external monitors see them within 90 seconds; what do you do first?

    deploymentmonitoringrouting
  • 57

    A partner accidentally originates your /22 from AS 64580, attracting 35% of traffic despite your valid /22 announcement; how do you recover safely?

  • 58

    Users in 2 of 7 regions receive SERVFAIL for api.example.com after a delegation change, while direct queries to 3 authoritative servers work; how do you isolate it?

    delegationapiqueries
  • 59

    A DNSSEC key rollover causes 38% validation failures, and the zone's DNSKEY TTL is 900 seconds; what recovery sequence do you choose?

    validation
  • 60

    One of 8 DNS anycast PoPs returns stale records to 12% of clients, but its BGP announcement and health check remain green; how do you contain it?

    dnsroutingnetwork-services
  • 61

    Checkout RTT rises from 28 to 96 ms for 1 region after 02:10, but interface counters show 0 errors; how do you prove where the path changed?

    types
  • 62

    TCP uploads lose 3% of packets in one direction across a 4-hop WAN, while downloads are clean and MTR disagrees by hop; how do you localize it?

    networkingprotocolsconflict
  • 63

    After an ACL change, 14 hosts in VLAN 210 can reach a payroll subnet that should allow only 3 jump servers; how do you stop and explain the leak?

    networkingip-addressingswitching
  • 64

    Tenant A learns 620 Tenant B prefixes after an EVPN policy edit, although the VRFs use different VNIs; how do you contain the route leak?

    routing
  • 65

    An uplink fails, but HSRP leaves the VIP on the isolated active switch and 4,000 clients blackhole for 6 minutes; what do you fix?

    switchinghigh-availability
  • 66

    A VRRP pair fails over in 1 second, but 30% of sessions reset because return traffic still uses the old firewall path; how do you recover?

    sessionsnetwork-securitynetworking
  • 67

    OSPF neighbors on 12 links oscillate between EXSTART and FULL after a software upgrade, causing 9-second outages; how do you diagnose it?

    routing
  • 68

    A misconfigured redistribution creates 18,000 OSPF external LSAs and drives 40 routers above 85% CPU; what is your recovery order?

    routingredis
  • 69

    An IS-IS Level 2 router reports overload and 6 sites take a 45 ms detour after a memory alarm; how do you handle the event?

    soft-skillsmemoryrouting
  • 70

    A new trunk creates an STP loop, broadcast traffic reaches 18 Gbps, and 22 access switches hit 95% CPU; how do you restore the campus?

    switching
  • 71

    A 4-member LACP bundle shows all links up, but 25% of flows fail after one switch reloads; what specific fault do you suspect and test?

    switching
  • 72

    A 100 Gbps optic accumulates 900 CRC errors per minute only above 70% load, and FEC corrected errors rise 12x; how do you decide what to replace?

  • 73

    One EVPN tenant loses 240 endpoints after a template change, while 19 other tenants remain healthy; how do you bound and recover it?

    endpoints
  • 74

    A virtual machine triggers 46 EVPN MAC mobility events in 60 seconds between 2 leaves, causing intermittent 8-second outages; what do you do?

  • 75

    BUM traffic in VNI 50120 jumps from 200 Mbps to 14 Gbps and saturates 3 VTEPs after an endpoint rollout; how do you contain it?

    endpoints
  • 76

    Jumbo storage traffic fails on 2 of 8 VXLAN ECMP paths after a spine replacement, while 1,500-byte traffic is clean; what is your recovery plan?

    recoverydatacenter
  • 77

    After adding a 3rd internet exit, 18% of TCP sessions hit firewall out-of-state drops although both directions are reachable; how do you restore service?

    networkingprotocolssessions
  • 78

    A NAT gateway reaches 98% of its 64,000 ports for 1 public IP, and API connection failures climb to 17%; what do you change first?

    gatewaynetworkingnetwork-services
  • 79

    An L4 load balancer loses 2 of 12 nodes and resets 9% of 1.6 million TCP sessions despite healthy backend servers; how do you limit damage?

    load-balancingnetworkingprotocols
  • 80

    An L7 proxy deployment raises 5xx responses from 0.2% to 11% because health checks mark 40% of backends down; how do you decide rollback?

    deploymenthealth-checksrollback
  • 81

    A 620 Gbps UDP flood saturates 2 x 100 Gbps internet links within 3 minutes; what actions do you take before local ACLs can help?

    networkingprotocols
  • 82

    A 4 Mpps SYN flood drives an edge router to 92% CPU while links remain below 20 Gbps; how do you protect the control plane?

    routing
  • 83

    A campus DHCP scope has 0 free addresses for 6,200 clients, while the standby server reports PARTNER-DOWN for 18 minutes; how do you recover leases safely?

    network-services
  • 84

    After a switch upgrade, 1,800 IPv6 clients install a rogue default router with 30-second lifetime and lose SaaS access every 5 minutes; how do you stop it?

    ip-addressingroutingswitching
  • 85

    A 10 Gbps Direct Connect drops at 14:05, but VPN fallback delivers only 2 Gbps and 3 critical applications need 1.4 Gbps; how do you fail over?

    network-security
  • 86

    A Transit Gateway route-table edit sends 27 VPCs to a missing inspection attachment, raising timeout errors to 23%; how do you restore connectivity?

    gatewayresiliencerouting
  • 87

    An SD-WAN policy oscillates 36 branches between MPLS and DIA every 90 seconds when loss hovers near 1%; how do you stabilize service?

    wan
  • 88

    An MPLS VPN loses 8 sites because an LDP label is missing although the IGP route remains present; how do you isolate the forwarding break?

    routingwannetwork-security
  • 89

    On a 500 Mbps failed WAN path, voice uses 95 Mbps although its LLQ is configured for 15%, starving a 150 Mbps business class; what do you change?

    config
  • 90

    Automation pushes a wrong prefix list to 160 routers in 4 minutes, withdrawing 12,000 customer routes; how do you stop the rollout?

    routing
  • 91

    A configuration job succeeds on 73 of 100 switches, then loses API access and leaves mixed VLAN state; do you roll forward or back?

    switchingapiconfig
  • 92

    During a 4,000-user HSRP blackhole, a junior engineer wants to clear ARP on all access switches; how do you mentor them while restoring service?

    protocolsswitchinghigh-availability
  • 93

    A mid-level engineer sees 7 BGP resets in 10 minutes and proposes global dampening; how do you coach the investigation and containment?

    routing
  • 94

    A junior engineer replaces 2 optics and a fiber after seeing 300 CRC errors per minute, but the fault remains; how do you redirect the next 45 minutes?

    react
  • 95

    A 47-minute packet-loss incident ended after an ECMP member was drained, but no device alarm fired despite 6% checkout failures; what belongs in the postmortem?

    incidents
  • 96

    A DNS incident is mitigated in 8 minutes, but 9% of clients still fail for 1 hour because resolvers cached NXDOMAIN; how do you judge recovery?

    dnsnetwork-servicescaching
  • 97

    A route-policy change cuts latency from 70 to 35 ms but causes 2% loss on 1 carrier; do you roll back after 12 minutes or tune forward?

    routinglatencyrollback
  • 98

    IPsec tunnels at 60 branches drop every 55 minutes after a key-policy change, interrupting voice for 20 seconds; how do you isolate rekey failure?

    network-security
  • 99

    A telemetry change sends 180,000 packets per second to router CPUs, causing 24 OSPF adjacencies to flap; how do you recover the control plane?

    routing
  • 100

    At 09:00, 11 sites lose all applications and 3 recent network changes are plausible; how do you run the first 30 minutes and choose recovery?