The multi-homing lab
ended with a node that has two independent routed uplinks, a BGP host
route over each, and maximum-paths 8 on every speaker. The kernel
route table agrees:
$ ip -6 route show fd00:d::9
fd00:d::9 nhid 22 proto bgp metric 20 pref medium
nexthop via fe80::a8c1:abff:fe6b:6dd5 dev enp1s3 weight 1
nexthop via fe80::a8c1:abff:feaa:5b76 dev ens2 weight 1
Two next hops, weight 1 each. Then you send traffic.
The problem, in one measurement
121 UDP flows from a pod, each with a distinct source port, one
destination behind both uplinks. tcpdump on each uplink counts flows:
ens2 (tor1): 121
enp1s3 (tor2): 0
100% / 0%. Every flow takes the same link. Setting
net.ipv6.fib_multipath_hash_policy=1 — the L4 hash policy — and
repeating: still 121 / 0. The sysctl is not the missing piece.
Why: the lookup carries no flow
Cilium’s native fast path asks the kernel for the next hop with the bpf_fib_lookup() helper and
redirects to the returned interface itself. The helper does run the
multipath selection — but it hashes only what the caller puts into
struct bpf_fib_lookup, and Cilium fills in family, ifindex, source and
destination address.
So for one destination, every flow presents an identical hash key. The hash is deterministic; the “choice” is always the same next hop. Each lookup returns a single next hop — spreading is what you get in aggregate over many flows, and only if the keys differ between flows.
| |
flowinfo feeds the default (L3) IPv6 hash policy, which hashes
source, destination and flow label. The ports feed policy 1. A quick
Cilium-free check on an isolated netns (ip route get takes the same
skb-less path as the helper) confirms the kernel side works: with
sport/dport varied across 96 lookups, policy 0 pins 96/0, policy 1
spreads 55/41. The kernel was never the problem; the lookup was simply
asked an L3-only question.
The change
On my cilium fork (main,
four commits): fill the lookup input with the
packet’s flow keys before every forwarding bpf_fib_lookup. Lookup
input only — the packet is never modified.
| |
The decisions behind it, briefly:
- Synthesise a flow label when the packet has none. The kernel auto-generates labels for TCP sockets, but UDP typically leaves the label zero — and the label is what the default hash policy uses. Deriving one from the ports (only the 20 label bits, never written to the packet) means IPv6 spreads with no sysctl at all.
- TCP, UDP and SCTP only — the same set Cilium’s conntrack tracks. QUIC is UDP as far as a forwarder is concerned; the 4-tuple already carries it. IPv6 extension headers and non-first IPv4 fragments fall outside the gate and are left alone.
- The caller passes what it knows. The helpers take the header
pointer and offset from the call site instead of re-parsing packets
inside generic code. The nodeport tails that pre-build their own
fib_paramscall the same helper; the two rev-DNAT paths whose params describe a synthetic tunnel outer while the packet may still be the inner one are deliberately skipped — mixing layers in one hash key is how you get silently wrong routing. - Tunnel underlay spreads too. Where the packet reaching the lookup is the encapsulated outer, the outer UDP source port (which Cilium already derives from the inner flow) enters the hash, so overlay traffic between two nodes stops pinning to one uplink as well.
- IPv4 has no flow label, so it gets the 5-tuple only and needs
net.ipv4.fib_multipath_hash_policy=1on the node. - Egress-gateway (forced oif, pre-SNAT ports) and the source-address
lookups (
fib_lookup_src_*resolve an address, not a next hop) are untouched.
Files: bpf/lib/fib.h (+97), bpf/lib/l4.h (+7),
bpf/lib/nodeport.h (+10), bpf/lib/nodeport_egress.h (+15), and the
tests bpf/tests/fib_tests.c (+326),
bpf/tests/tc_nodeport_lb4_dsr_backend.c (+23) and
bpf/tests/tc_nodeport_lb6_dsr_backend.c (+32) — pktgen unit tests that
assert the recorded lookup keys: label kept for TCP, label synthesised
for UDP, ICMPv6 untouched, later fragments skipped.
One kernel surprise
The lab node image ran Ubuntu’s -kvm kernel flavour, and IPv4 ECMP
simply does not exist there:
$ grep IP_ROUTE_MULTIPATH /boot/config-5.15.0-1103-kvm
# CONFIG_IP_ROUTE_MULTIPATH is not set
Not a kernel-version issue — a flavour choice in the minimized KVM
config. The lab image now installs linux-image-virtual
(5.15.0-187-generic, CONFIG_IP_ROUTE_MULTIPATH=y), and with it the
IPv4 BGP route materialises as real multipath, RFC 8950 next hops and
all:
$ ip -4 route show 10.9.9.9
10.9.9.9 nhid 25 proto bgp src 10.2.0.104 metric 20
nexthop via inet6 fe80::a8c1:abff:feaa:5b76 dev ens2 weight 1
nexthop via inet6 fe80::a8c1:abff:fe6b:6dd5 dev enp1s3 weight 1
After
160 random-source-port UDP flows per run, counted on both uplinks:
| AF | hash policy | ens2 | enp1s3 | split | |
|---|---|---|---|---|---|
| A | v6 | 0 (default) | 88 | 72 | 55/45 — spreads on the default policy |
| B | v6 | 1 (L4) | 89 | 71 | 55/44 |
| C | v4 | 0 (default) | 160 | 0 | 100/0 — expected; v4 L3 hash has no label |
| D | v4 | 1 (L4) | 85 | 75 | 53/46 |
Row A is the point of the exercise: IPv6 spreading with zero node configuration, because every flow now carries a label into the lookup. Row C is the honest control — IPv4’s L3 hash has nothing per-flow to hash for a single destination, so the sysctl (row D) is part of the IPv4 answer.
Same pod source address, different source ports, both uplinks:
=== ens2 (tor1) ===
IP6 fd02:b::1a.16319 > fd00:d::9.9999: UDP, length 1
IP6 fd02:b::1a.8823 > fd00:d::9.9999: UDP, length 1
=== enp1s3 (tor2) ===
IP6 fd02:b::1a.21066 > fd00:d::9.9999: UDP, length 1
IP6 fd02:b::1a.15287 > fd00:d::9.9999: UDP, length 1
Interface tx_packets deltas agree with the pcaps (155/165 over 300
flows), so it is not a capture artefact.
Running it
The code lives on the fork’s main, on top of upstream cilium/cilium
main (rebased, not diverged — the seven touched files carry the whole
change):
| |
Both images matter: an agent built from main waits for CRDs that only
an operator built from the same tree registers. Mixing a new agent
with an older operator leaves the agent politely waiting forever.
Ship them to the nodes (containerd):
| |
Install with the in-tree chart and the same values as the multi-homing post, overriding the images:
| |
Then, per address family: IPv6 needs nothing. IPv4 needs the L4 hash policy and a kernel that has multipath at all:
| |
Notes
- Each
bpf_fib_lookupstill returns exactly one next hop. Nothing round-robins; a flow keeps its path (no reordering), and distinct flows land on distinct paths because their keys differ. - On a single-path route the extra keys change nothing — same lookup, same answer — so this is safe to run everywhere, not only on multi-homed nodes.
- The synthesized label exists only in the lookup key. If you run
flow-label-based
ip -6 rulepolicy routing, be aware the lookup now sees a label where the packet has none. - Lab, configs and measurement scripts: github.com/denizaydin/cilium-containerlab-bgp-multihoming — the code: github.com/denizaydin/cilium.
The content of this post was created by the author; AI was used for coding, editing and also restructuring the post