The multi-homing lab ended with a node that has two independent routed uplinks, a BGP host route over each, and maximum-paths 8 on every speaker. The kernel route table agrees:

$ ip -6 route show fd00:d::9
fd00:d::9 nhid 22 proto bgp metric 20 pref medium
 nexthop via fe80::a8c1:abff:fe6b:6dd5 dev enp1s3 weight 1
 nexthop via fe80::a8c1:abff:feaa:5b76 dev ens2 weight 1

Two next hops, weight 1 each. Then you send traffic.

The problem, in one measurement

121 UDP flows from a pod, each with a distinct source port, one destination behind both uplinks. tcpdump on each uplink counts flows:

ens2   (tor1): 121
enp1s3 (tor2): 0

100% / 0%. Every flow takes the same link. Setting net.ipv6.fib_multipath_hash_policy=1 — the L4 hash policy — and repeating: still 121 / 0. The sysctl is not the missing piece.

Why: the lookup carries no flow

Cilium’s native fast path asks the kernel for the next hop with the bpf_fib_lookup() helper and redirects to the returned interface itself. The helper does run the multipath selection — but it hashes only what the caller puts into struct bpf_fib_lookup, and Cilium fills in family, ifindex, source and destination address.

So for one destination, every flow presents an identical hash key. The hash is deterministic; the “choice” is always the same next hop. Each lookup returns a single next hop — spreading is what you get in aggregate over many flows, and only if the keys differ between flows.

1
2
3
4
5
6
7
8
/* set if lookup is to consider L4 data - e.g., FIB rules */
__u8 l4_protocol;
__be16 sport;
__be16 dport;
...
union {
 /* inputs to lookup */
 __be32 flowinfo; /* AF_INET6, flow_label + priority */

flowinfo feeds the default (L3) IPv6 hash policy, which hashes source, destination and flow label. The ports feed policy 1. A quick Cilium-free check on an isolated netns (ip route get takes the same skb-less path as the helper) confirms the kernel side works: with sport/dport varied across 96 lookups, policy 0 pins 96/0, policy 1 spreads 55/41. The kernel was never the problem; the lookup was simply asked an L3-only question.

The change

On my cilium fork (main, four commits): fill the lookup input with the packet’s flow keys before every forwarding bpf_fib_lookup. Lookup input only — the packet is never modified.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
static __always_inline void
fib_params_set_l4_v6(struct bpf_fib_lookup_padded *fib_params,
       struct __ctx_buff *ctx, const struct ipv6hdr *ip6,
       int l3_off)
{
 __be32 flowinfo = *(const __be32 *)ip6 & IPV6_FLOWLABEL_MASK;
 __be16 ports[2];

 if (l4_proto_has_ports(ip6->nexthdr) &&
     l4_load_ports(ctx, l3_off + sizeof(struct ipv6hdr), ports) == 0) {
  fib_params->l.l4_protocol = ip6->nexthdr;
  fib_params->l.sport = ports[0];
  fib_params->l.dport = ports[1];

  if (!flowinfo)
   flowinfo = bpf_htonl(jhash_2words(ports[0], ports[1],
         JHASH_INITVAL)) &
       IPV6_FLOWLABEL_MASK;
 }

 fib_params->l.flowinfo = flowinfo;
}

The decisions behind it, briefly:

  • Synthesise a flow label when the packet has none. The kernel auto-generates labels for TCP sockets, but UDP typically leaves the label zero — and the label is what the default hash policy uses. Deriving one from the ports (only the 20 label bits, never written to the packet) means IPv6 spreads with no sysctl at all.
  • TCP, UDP and SCTP only — the same set Cilium’s conntrack tracks. QUIC is UDP as far as a forwarder is concerned; the 4-tuple already carries it. IPv6 extension headers and non-first IPv4 fragments fall outside the gate and are left alone.
  • The caller passes what it knows. The helpers take the header pointer and offset from the call site instead of re-parsing packets inside generic code. The nodeport tails that pre-build their own fib_params call the same helper; the two rev-DNAT paths whose params describe a synthetic tunnel outer while the packet may still be the inner one are deliberately skipped — mixing layers in one hash key is how you get silently wrong routing.
  • Tunnel underlay spreads too. Where the packet reaching the lookup is the encapsulated outer, the outer UDP source port (which Cilium already derives from the inner flow) enters the hash, so overlay traffic between two nodes stops pinning to one uplink as well.
  • IPv4 has no flow label, so it gets the 5-tuple only and needs net.ipv4.fib_multipath_hash_policy=1 on the node.
  • Egress-gateway (forced oif, pre-SNAT ports) and the source-address lookups (fib_lookup_src_* resolve an address, not a next hop) are untouched.

Files: bpf/lib/fib.h (+97), bpf/lib/l4.h (+7), bpf/lib/nodeport.h (+10), bpf/lib/nodeport_egress.h (+15), and the tests bpf/tests/fib_tests.c (+326), bpf/tests/tc_nodeport_lb4_dsr_backend.c (+23) and bpf/tests/tc_nodeport_lb6_dsr_backend.c (+32) — pktgen unit tests that assert the recorded lookup keys: label kept for TCP, label synthesised for UDP, ICMPv6 untouched, later fragments skipped.

One kernel surprise

The lab node image ran Ubuntu’s -kvm kernel flavour, and IPv4 ECMP simply does not exist there:

$ grep IP_ROUTE_MULTIPATH /boot/config-5.15.0-1103-kvm
# CONFIG_IP_ROUTE_MULTIPATH is not set

Not a kernel-version issue — a flavour choice in the minimized KVM config. The lab image now installs linux-image-virtual (5.15.0-187-generic, CONFIG_IP_ROUTE_MULTIPATH=y), and with it the IPv4 BGP route materialises as real multipath, RFC 8950 next hops and all:

$ ip -4 route show 10.9.9.9
10.9.9.9 nhid 25 proto bgp src 10.2.0.104 metric 20
 nexthop via inet6 fe80::a8c1:abff:feaa:5b76 dev ens2 weight 1
 nexthop via inet6 fe80::a8c1:abff:fe6b:6dd5 dev enp1s3 weight 1

After

160 random-source-port UDP flows per run, counted on both uplinks:

AFhash policyens2enp1s3split
Av60 (default)887255/45 — spreads on the default policy
Bv61 (L4)897155/44
Cv40 (default)1600100/0 — expected; v4 L3 hash has no label
Dv41 (L4)857553/46

Row A is the point of the exercise: IPv6 spreading with zero node configuration, because every flow now carries a label into the lookup. Row C is the honest control — IPv4’s L3 hash has nothing per-flow to hash for a single destination, so the sysctl (row D) is part of the IPv4 answer.

Same pod source address, different source ports, both uplinks:

=== ens2 (tor1) ===
IP6 fd02:b::1a.16319 > fd00:d::9.9999: UDP, length 1
IP6 fd02:b::1a.8823  > fd00:d::9.9999: UDP, length 1
=== enp1s3 (tor2) ===
IP6 fd02:b::1a.21066 > fd00:d::9.9999: UDP, length 1
IP6 fd02:b::1a.15287 > fd00:d::9.9999: UDP, length 1

Interface tx_packets deltas agree with the pcaps (155/165 over 300 flows), so it is not a capture artefact.

Running it

The code lives on the fork’s main, on top of upstream cilium/cilium main (rebased, not diverged — the seven touched files carry the whole change):

1
2
3
git clone https://github.com/denizaydin/cilium && cd cilium
make docker-cilium-image DOCKER_IMAGE_TAG=flowkeys
make docker-operator-generic-image DOCKER_IMAGE_TAG=flowkeys

Both images matter: an agent built from main waits for CRDs that only an operator built from the same tree registers. Mixing a new agent with an older operator leaves the agent politely waiting forever.

Ship them to the nodes (containerd):

1
2
docker save quay.io/cilium/cilium:flowkeys | ctr -n k8s.io images import -
docker save quay.io/cilium/operator-generic:flowkeys | ctr -n k8s.io images import -

Install with the in-tree chart and the same values as the multi-homing post, overriding the images:

1
2
3
4
5
6
helm upgrade --install cilium ./install/kubernetes/cilium -n kube-system \
  -f cilium-values.yaml \
  --set image.override=quay.io/cilium/cilium:flowkeys \
  --set image.pullPolicy=IfNotPresent \
  --set operator.image.override=quay.io/cilium/operator-generic:flowkeys \
  --set operator.image.pullPolicy=IfNotPresent

Then, per address family: IPv6 needs nothing. IPv4 needs the L4 hash policy and a kernel that has multipath at all:

1
2
sysctl -w net.ipv4.fib_multipath_hash_policy=1
grep CONFIG_IP_ROUTE_MULTIPATH /boot/config-$(uname -r)   # must be =y

Notes

  • Each bpf_fib_lookup still returns exactly one next hop. Nothing round-robins; a flow keeps its path (no reordering), and distinct flows land on distinct paths because their keys differ.
  • On a single-path route the extra keys change nothing — same lookup, same answer — so this is safe to run everywhere, not only on multi-homed nodes.
  • The synthesized label exists only in the lookup key. If you run flow-label-based ip -6 rule policy routing, be aware the lookup now sees a label where the packet has none.
  • Lab, configs and measurement scripts: github.com/denizaydin/cilium-containerlab-bgp-multihoming — the code: github.com/denizaydin/cilium.

The content of this post was created by the author; AI was used for coding, editing and also restructuring the post