VXLAN EVPN separates a routable IP fabric underlay from an overlay that carries tenant Layer 2 and Layer 3 reachability between VXLAN Tunnel Endpoints (VTEPs). Current Cisco Nexus 9000 guidance describes the underlay as the network that advertises VTEP and BGP-peering reachability, while the overlay uses MP-BGP EVPN to distribute MAC, IP, prefix, and multihoming information. The two layers must be designed independently enough to troubleshoot, but coherently enough that overlay paths never depend on underlay reachability that is unstable or ambiguous.
Within Cisco Network Engineering, VXLAN is not a replacement for routing discipline. It adds encapsulation and an EVPN control plane on top of a leaf-spine IP network.
The existing VXLAN EVPN in the Data Center article covers the architectural trade-offs. This page focuses on underlay/overlay boundaries and operations.
The underlay exists to make VTEPs reachable
Cisco’s current Nexus design guide states that the primary purpose of the underlay is advertising reachability to VTEP loopbacks and BGP peering addresses.
Fast convergence, simplicity, and clean node bring-up behavior matter more than carrying tenant routes.
Keep tenant prefixes out of the underlay wherever possible so a routing problem can be isolated to fabric infrastructure rather than mixed with application segmentation.
Use loopbacks as stable VTEP and control-plane endpoints
Leaf switches typically use loopback addresses for VTEP/NVE source interfaces and BGP/EVPN peering.
The underlay must advertise those loopbacks through every spine path.
Stable loopbacks decouple tunnel/control-plane identity from physical interfaces and support ECMP across several routed links.
OSPF, IS-IS, and eBGP are current underlay options
Cisco Nexus 9000 documentation supports OSPF, IS-IS, and eBGP underlays depending on design.
The current design guide recommends an IGP underlay with an iBGP EVPN overlay as a preferred pattern, while eBGP/eBGP fabrics are also supported in relevant designs.
Choose one protocol model based on operational familiarity, scale, policy needs, and automation rather than mixing protocols unnecessarily inside one fabric.
The overlay control plane is MP-BGP EVPN
EVPN distributes overlay reachability independently from the underlay IGP/BGP routes.
Leaf VTEPs advertise MAC/IP and prefix information; spines commonly act as route reflectors in iBGP overlay designs.
At least two route reflectors provide control-plane resiliency in a fabric with multiple spines according to Cisco’s design guidance.
Keep underlay and overlay troubleshooting separate
If two VTEPs cannot reach each other’s loopbacks, the overlay cannot work regardless of EVPN route state.
Verify physical links, MTU, IP addressing, IGP/eBGP adjacency, ECMP, loopback routing, and BFD/convergence before debugging VNIs or EVPN routes.
Only after underlay reachability is stable should operators inspect BGP EVPN sessions and route types.
MTU must account for VXLAN encapsulation
VXLAN adds encapsulation overhead to the original Ethernet frame.
Cisco Nexus guidance commonly recommends a jumbo MTU such as 9216 across VTEP-to-VTEP paths when servers may use 9000-byte frames.
An inconsistent MTU can create selective failures where small control traffic succeeds but large application packets fragment or drop.
VNI design is the overlay segmentation contract
Layer 2 VNIs extend broadcast domains between VTEPs; Layer 3 VNIs provide routed tenant/VRF connectivity.
Map VLANs, bridge domains, VRFs, and VNIs consistently through automation.
VNI numbering should be treated like an address plan with ownership and change control rather than chosen ad hoc on each leaf.
Route types should match the service being built
EVPN uses different route types for MAC/IP advertisement, inclusive multicast, IP prefixes, multihoming, and other functions.
Operators should know which route type proves the expected control-plane state for the failing service.
For example, an inter-subnet routing problem may require inspecting prefix/MAC-IP information rather than only checking that the BGP neighbor is established.
BUM handling is an explicit design choice
Broadcast, unknown unicast, and multicast traffic can be handled through multicast underlay mechanisms or ingress replication depending on design and scale.
Current Nexus releases add optimized multicast capabilities such as selective EVPN multicast route types for relevant deployments.
Choose a model that matches scale and operational skills; BUM behavior is part of the data plane, not a minor implementation detail.
ECMP should be expected throughout the underlay
Leaf-spine fabrics rely on equal-cost routed paths through several spines.
Validate hashing, link utilization, and failure convergence rather than assuming one deterministic physical path.
BGP Best-Path Decisions is relevant for understanding control-plane selection where BGP is used, but forwarding may still use ECMP over multiple equal paths.
VXLAN underlay/overlay succeeds when each layer has a clear job
The mature fabric keeps VTEP reachability simple and fast in the underlay, uses MP-BGP EVPN for tenant reachability, maintains consistent MTU/VNI/VRF design, provides route-reflector redundancy, and troubleshoots physical/routing health before overlay state.
VXLAN becomes manageable when the team can answer whether a failure is “the tunnel endpoints cannot reach each other” or “the overlay did not advertise the tenant information”—without mixing the two.
Underlay addressing should be simple enough to generate. Point-to-point leaf-spine links commonly use small subnets or unnumbered/structured addressing patterns depending on platform design; loopbacks come from dedicated ranges. Whatever convention is chosen, automation should be able to derive it deterministically and monitoring should know which addresses represent physical links, routing IDs, VTEPs, and service loopbacks.
Failure domains should be explicit. A single spine link failure should use ECMP around the problem, while a leaf failure affects only attached endpoints. Route-reflector placement, vPC/multihoming, border leafs, and external routing can introduce larger failure domains. Test each failure separately and confirm both underlay convergence and EVPN route withdrawal happen inside the application SLO.
BFD can accelerate failure detection for underlay routing where supported and justified, but aggressive timers consume control-plane resources and can react to transient congestion. Set timers from fabric scale and operational goals rather than copying the fastest values from a lab. Convergence should be measured under link, node, and control-plane failures.
Anycast gateway designs keep the same default-gateway IP/MAC available on multiple VTEPs so endpoints can route locally at the leaf. This reduces hairpinning and supports mobility across VXLAN segments. The gateway configuration must be consistent across all participating leafs, and troubleshooting should verify both local SVI/VRF state and EVPN advertisements.
ARP/ND suppression can reduce broadcast traffic by using EVPN-learned endpoint information, but stale control-plane entries can create confusing reachability problems. When a host cannot resolve a neighbor, inspect endpoint learning, EVPN MAC/IP routes, suppression state, and host mobility rather than immediately blaming the underlay.
Border leafs form the boundary between fabric tenant routes and external networks. Keep route leaking, default routes, VRF handoff, firewall/service insertion, and internet/DCI policy explicit. A well-designed internal fabric can still fail operationally if border leaves advertise the wrong tenant prefixes or create asymmetric paths through external security devices.
vPC or EVPN multihoming lets dual-attached servers/appliances survive a leaf failure. The design must coordinate underlay reachability, EVPN Ethernet Segment/multihoming state, loop prevention, and host LAG behavior. Test orphan ports and one-sided failures, not just clean shutdown of a whole leaf.
Fabric MTU should be verified end to end with probes that exercise the real overlay path. Interface configuration showing 9216 everywhere is not proof when intermediate port-channels, firewalls, DCI links, or virtual switches use smaller MTUs. Use DF-bit/payload testing and device counters to find the exact hop dropping large encapsulated frames.
Operations should maintain separate dashboards for underlay adjacencies/ECMP/loopbacks and overlay BGP EVPN/VNI/endpoint state. This prevents an EVPN route-table problem from being hidden behind a generic ‘fabric down’ alert and lets the on-call engineer pick the right troubleshooting branch quickly.
Overlay policy should account for tenant route leaking and shared services explicitly. Firewalls, DNS, load balancers, monitoring, and internet gateways often sit in shared VRFs that multiple tenants must reach. Use controlled route leaking and service-insertion design instead of ad hoc static routes that bypass the EVPN segmentation model.
Fabric automation should validate intent before deployment. Generate underlay addressing, routing, VNI/VLAN/VRF mappings, route-reflector neighbors, anycast gateways, and border policy from a source of truth, then compare device state with intent. VXLAN scale makes hand-edited inconsistencies difficult to see until one leaf carries a mismatched VNI or missing RT.
Telemetry should include NVE peer state, BGP EVPN prefixes by route type, endpoint moves, ARP/ND suppression counters, underlay adjacency/BFD status, interface errors, and ECMP utilization. Baselines help distinguish one bad host or VTEP from a control-plane-wide problem during incident response.
Disaster recovery and DCI should preserve the underlay/overlay boundary. Extending VNIs between sites, using EVPN multisite, or leaking tenant routes into WAN/DCI adds another control plane. Keep site-local failures from propagating globally and document which prefixes/MACs are intentionally stretched versus summarized at the inter-site boundary.
Change windows should validate control-plane scale as well as reachability. Adding hundreds of VNIs or tenants can increase BGP EVPN state, ARP/ND tables, and route-reflector memory even when link utilization is low. Capacity reviews should include route counts and convergence time so growth does not create a control-plane bottleneck.
Fabric documentation should make the encapsulation path visible from endpoint to endpoint: source VLAN/VRF, VNI, source VTEP, underlay next hops, destination VTEP, and egress VLAN/VRF. This simple mapping shortens incident response because teams can test each boundary systematically.
Troubleshooting should prove the underlay first: IP reachability, MTU, routing, and redundancy. Only then should operators interpret VXLAN control-plane and overlay state, because an overlay symptom can be the visible consequence of a transport problem below it.