In 20 years, you will be more dissapointed by what you didn't do than by what you did.

LACP Load Balancing Troubleshooting: Why One Flow Uses Only One Link

A four-member port channel can be healthy while one transfer uses only one member. Before replacing optics or increasing buffers, separate aggregate capacity from the forwarding decision for an individual flow. Linux bonding documentation explicitly notes that its layer3+4 policy can distribute traffic to one peer across members, but a single connection does not span multiple members.[1]

This guide is a practical diagnostic workflow for conventional Ethernet LACP and Linux 802.3ad bonds—not a claim about RDMA multipathing or GPU fabric performance. The design and test cases below are hypothetical; no production benchmark or failover test was run for this article.

The problem: all links are up, but the application is slow

Start with a precise symptom: one backup stream is slow, many clients are slow, or only traffic in one direction is slow. Those are different investigations. Do not use the port-channel interface's advertised aggregate speed as the expected throughput of every application.

Linux 802.3ad bonding groups links into aggregators, and its transmit hash policy selects the member used for outgoing traffic.[1] LACP health and useful traffic distribution should therefore be separate acceptance checks: first establish that the intended members participate, then measure which members carry the workload.

For related fabric troubleshooting, keep the Data Center hub nearby. If only large packets fail, follow the separate EVPN/VXLAN MTU workflow rather than assuming a hashing problem.

A concrete sizing example: four 100 Gb/s members

Hypothetical design: a Linux data-transfer host has four 100 Gb/s Ethernet members in one LACP bond to a compatible switch. Assume the rest of the path and the remote test system can support the aggregate load; this assumption must be checked in a real deployment.

Tool-calculated theoretical capacity, in one direction:

  • Four members: 4 × 100 = 400 Gb/s, or 50 GB/s in decimal units.
  • One member: 100 ÷ 8 = 12.5 GB/s.
  • These are nominal link-rate conversions, not TCP payload rates, storage throughput, or benchmark results.

For the Linux layer3+4 behavior documented above, a single TCP connection remains on one member.[1] In this example its link-rate ceiling is therefore 100 Gb/s, not 400 Gb/s, before other bottlenecks and overhead. Multiple connections create opportunities for distribution, not a guarantee that each connection lands on a different member.

Single connection on one LACP member and multiple connections distributed across members
Source: original Network freak diagram, based on Linux bonding documentation [1]. Hypothetical topology; green links illustrate traffic placement, not measured throughput.

Diagnose the two transmit directions separately

Treat the host-to-switch and switch-to-host directions as separate measurement cases. Record the host's transmit policy and the switch's documented egress port-channel policy. Do not assume that changing the host's setting changes the switch's forwarding decision.

The Linux bonding default transmit policy is layer2; its documentation says that traffic to a particular network peer uses the same member.[1] Layer2+3 adds IP information but still places traffic to a particular network peer on the same member, while layer3+4 can distinguish connections through transport-layer information.[1]

That distinction matters when choosing tests. Opening more TCP connections between the same endpoints does not exercise new hash inputs if the active policy ignores transport ports. Instead of immediately editing configuration, write down which fields each transmitting device actually uses and choose tests that vary those fields.

Read-only Linux evidence collection

Unexecuted example commands: adapt the bond and physical interface names to your lab. These examples inspect state; they do not configure a bond.

ip -d link show dev bond0
ip -s link show dev bond0
ip -s link show dev eth1
ip -s link show dev eth2
ip -s link show dev eth3
ip -s link show dev eth4
python3 -c 'from pathlib import Path; print(Path("/proc/net/bonding/bond0").read_text())'

The bonding driver exposes bond information through /proc/net/bonding, and its documented settings include mode, link monitoring, LACP rate and transmit hash policy.[1] Save the complete output privately. Compare member link status, aggregator membership and partner information with the switch's LACP detail view. Remove real hostnames, addresses and hardware identifiers before sharing evidence externally.

On the switch, use the platform-specific equivalents of port-channel summary, LACP neighbor/detail, per-member interface counters and load-balance policy display. Exact commands differ by platform and release; do not paste syntax from an unrelated operating system.

Build a controlled throughput test matrix

iperf3 supports a server/client model, parallel streams with -P, reverse-direction testing with -R, test duration with -t and JSON output with -J.[2] Check the installed manual before use: the project notes that its web-rendered manual can differ from the installed version.[2]

Unexecuted lab examples: replace TEST_SERVER with an authorized test endpoint. Permit access only between the intended lab hosts, and avoid running saturation tests on a shared production path without an approved window.

# On the authorized test server
iperf3 -s

# On the client: run separately, not concurrently
iperf3 -c TEST_SERVER -t 30 -P 1 -J > single-forward.json
iperf3 -c TEST_SERVER -t 30 -P 8 -J > parallel-forward.json
iperf3 -c TEST_SERVER -t 30 -P 1 -R -J > single-reverse.json
iperf3 -c TEST_SERVER -t 30 -P 8 -R -J > parallel-reverse.json

Collect physical-member counters immediately before and after each test. Use counter deltas over the same interval, rather than comparing lifetime totals. Also record CPU utilization, retransmissions, interface errors and the remote endpoint's capacity. A throughput change alone does not identify the cause.

Observation Next check—not a definitive diagnosis
One stream uses one member; parallel streams spread Compare application connection count with the tested workload
Parallel streams still use one member Inspect hash fields, aggregator membership and endpoint diversity
Forward distributes; reverse does not Inspect the transmitting switch's egress policy and reverse-path bottlenecks
Members distribute but aggregate throughput stays low Check CPU, remote NICs, path capacity and test-tool limitations
Errors or drops rise on one member Investigate that member before changing the hash policy

If the application is storage-backed, repeat with its real data path after the network-only test. Record results separately. A successful synthetic network test is not acceptance evidence for disk performance, encryption throughput or the application's concurrency model.

When a hash change is justified—and when it is not

Consider a policy change only after the evidence shows insufficient hash diversity and the platform documentation supports the proposed setting. Linux documents an important caveat: layer3+4 is not fully 802.3ad compliant because a conversation containing both fragmented and unfragmented traffic can be mapped differently, potentially producing out-of-order delivery.[1] It is not a universal copy-and-paste performance fix.

For tunneled workloads, Linux also documents encap2+3 and encap3+4 policies that may use inner headers through the flow dissector.[1] Verify actual encapsulation support and traffic distribution in your kernel and switch combination; do not infer it solely from the policy name.

If the real requirement is one very large transfer, compare three alternatives: a faster individual Ethernet link, application-supported concurrent connections, or an explicitly supported multipath design. Choose based on the application's behavior and failure requirements, not just the sum of NIC port speeds. Do not replace LACP with packet striping as an improvised production workaround.

Failure domains and deployment acceptance

For the hypothetical single-switch design, multiple host links do not remove the switch as a common dependency. If switch-level availability is required, evaluate the supported multi-chassis design separately; do not simply spread an ordinary LAG across unrelated switches.

Use this proposed acceptance checklist before closing a change:

  • Confirm all intended members are in the expected active aggregator on both ends.
  • Save the original and proposed hash policies, affected interfaces and rollback method.
  • Record single-stream and parallel-stream results in both directions, with per-member deltas.
  • Agree application-specific throughput and interruption limits before the test.
  • During a maintenance window, remove one approved member and measure actual application recovery; restore it and verify it rejoins correctly.
  • Test the switch failure domain separately if the architecture claims switch redundancy.
  • Stop and roll back for unexplained session loss, rising errors or throughput regression.

Keep rollback access independent of the tested bond. The out-of-band management checklist is a useful companion for planning recovery.

Practical takeaway

A healthy LACP bundle is not a promise that one connection can consume every member. Establish the active aggregator, identify the hash inputs in each direction, then correlate controlled traffic with physical-member counters. Change configuration only when those measurements support it—and accept the result against the real application, not the aggregate port label.

Sources

  1. https://docs.kernel.org/networking/bonding.html — Linux Ethernet Bonding Driver HOWTO
  2. https://software.es.net/iperf/invoking.html — Invoking iperf3

Comments

0 Responses to "LACP Load Balancing Troubleshooting: Why One Flow Uses Only One Link"

Post a Comment

Popular Posts