I lost an afternoon to a “routing loop” that turned out not to be a loop at all. It was an EVPN all-active multihoming feature quietly not working in a virtual lab — and the packets ping-ponging between two switches were a symptom, not the disease.
This post is two things: a refresher on how Arista ESI (Ethernet Segment) multihoming actually forwards traffic, and a write-up of the specific trap you can fall into when you build EVPN designs on vEOS/containerlab instead of real hardware.
First, what ESI multihoming is
In an EVPN/VXLAN fabric, a host that needs redundancy is connected to two leaf switches at once with a single logical link (an MLAG-like LAG, but signalled through EVPN). That shared segment is identified by an Ethernet Segment Identifier (ESI). When both leafs forward for it at the same time, it’s all-active multihoming.
flowchart TB FW["DC Firewall<br/>(one logical host)"] FW ===|"LAG · one Ethernet Segment (ESI)"| L1["blsw1<br/>leaf / VTEP"] FW ===|"same ESI"| L2["blsw2<br/>leaf / VTEP"] L1 --- SP["spine / VXLAN fabric"] L2 --- SP SP --- R["remote leaf<br/>(another VTEP)"] R --- SRV["server"]
Making “two switches, one host, both active” work correctly takes four EVPN route types working together:
- Type-4 — Ethernet Segment route: the two leafs discover they share the same ESI and elect a Designated Forwarder (DF) for broadcast/unknown/multicast traffic.
- Type-1 — Ethernet A-D (auto-discovery) route: advertises “this ESI is reachable through me.” This is what enables aliasing — a remote VTEP can load-balance unicast to both leafs even if only one of them actually advertised the host’s MAC.
- Type-2 — MAC/IP route: advertises the host’s MAC (and IP). Because it carries the ESI, every remote VTEP understands “reachable via the segment,” i.e. via either leaf.
- Local bias / split-horizon: each leaf prefers its own local link to the host and filters echoes, so the host is reached directly and frames don’t loop.
The important mental model: the host’s address is reachable via the segment, and each leaf turns that into “use my own local port.” Remote VTEPs spray traffic at both leafs (aliasing); each leaf lands it on the host locally (local bias). No single leaf is a bottleneck, and nothing loops.
The setup that broke
In my multi-site lab the DC firewall is a host, dual-homed all-active to two border leafs
(blsw1 / blsw2) over an ESI LAG. The two leafs share an anycast /29 SVI (.10 on one, .11
on the other, .14 the virtual IP), and the firewall originates a default route that the rest of the
fabric follows.
So “reach the firewall” means: a remote device resolves the firewall’s next-hop over that shared
/29, and the fabric is supposed to deliver it to whichever leaf can reach the firewall locally.
The symptom: a .10 ↔ .11 ping-pong
A traceroute toward the firewall bounced between the two border leafs and never arrived:
sequenceDiagram participant S as server (remote) participant B1 as blsw1 (.10) participant B2 as blsw2 (.11) participant FW as firewall S->>B1: packet to firewall IP Note over B1: no working local adjacency —<br/>believes MAC is behind blsw2 B1->>B2: forward over peer / overlay Note over B2: believes MAC is behind blsw1 B2->>B1: forward back Note over B1,B2: ping-pong .10 ↔ .11 (looks like a loop)
It looks like a classic L2/L3 loop. It isn’t. Neither switch has a wrong static route or a misconfigured next-hop. What’s actually happening is that the all-active ESI forwarding never engaged: the leaf that received the packet didn’t have a working local adjacency to the firewall, so — lacking aliasing/local-bias — it shipped the packet to its peer, which was in exactly the same situation. Two “helpful” switches, forever handing the packet to each other.
The root cause: a virtual dataplane, not a design bug
Here’s the trap. vEOS-lab is a software dataplane. Its EVPN control plane speaks all the right
route types, but it does not implement the all-active ESI L3 forwarding path — the Type-1 aliasing
and Type-2 local-bias behaviour that hardware ASICs do natively. So reaching a host through the
shared subnet is ambiguous: both leafs advertise the /29, and nothing disambiguates which one can
actually deliver to the host. The result is the bounce.
On real Arista hardware this simply works — aliasing load-balances to both leafs, each has a live local link, and Type-2 MAC/IP + local bias make the host reachable via either. The design is fine; the emulator is the limitation.
Rule of thumb: if a “loop” only reproduces in vEOS/containerlab and vanishes on hardware, suspect a dataplane feature gap (ESI all-active, VXLAN bridging edge cases) before you rewrite your config.
The fix: hand-generate the host /32
If the fabric can’t figure out which leaf owns the host, we tell it explicitly. Two commands:
| |
ip attached-host route export makes each leaf install a /32 host route for the firewall — but
only on the leaf that has actually ARP-resolved it on its own local link. redistribute attached-host injects that /32 into BGP/EVPN. Now forwarding follows an explicit, leaf-specific host
route instead of the ambiguous /29:
flowchart LR FW["firewall<br/>10.32.228.x"] -->|"ARP resolved on the local link"| B1["blsw1<br/>ip attached-host route export"] B1 -->|"/32 host route<br/>redistribute attached-host"| NET["BGP / EVPN"] NET -->|"/32 is longest-prefix,<br/>beats the /29"| DST["blsw2 and remote leafs<br/>forward straight to blsw1"]
The /32 is longest-prefix, so it wins over the /29. Traffic goes directly to the leaf that can
actually deliver to the firewall — no bounce. And because the /32 only exists while ARP is resolved,
it also tracks liveness: if the firewall drops, the host route withdraws.
What you’ve really done here is hand-generate, in the control plane, the host route that working EVPN Type-2 + ESI aliasing would have produced for you on hardware. It’s a lab workaround, not a production requirement.
Takeaways
- All-active ESI multihoming = one host, two leafs, both forwarding. It needs Type-4 (DF), Type-1 (aliasing), Type-2 (MAC/IP), and local-bias working together.
- A ping-pong between two multihoming leafs is usually not a routing loop — it’s the all-active forwarding path failing to engage, so each leaf punts to the other.
- vEOS/containerlab emulates the EVPN control plane but not the all-active ESI L3 dataplane. Great for control-plane and topology testing; misleading for anything that depends on ASIC forwarding.
ip attached-host route export+redistribute attached-hostis a clean way to pin a multihomed host with an explicit /32 — invaluable in a virtual lab, unnecessary on real gear.
Build EVPN in a virtual lab to learn the protocol — just remember which parts of “it works on hardware” the emulator is quietly skipping.
The content of this post was created by the author; AI was used for editing and restructuring.