In 20 years, you will be more dissapointed by what you didn't do than by what you did.

RoCE PFC Buffer Sizing Guide: 400G Headroom Worksheet and ECN Checklist

A GPU cluster can pass a quiet link test and still collapse under synchronized traffic. Before buying switches with a bigger advertised buffer, ask a more precise question: how much traffic can arrive after a congested port requests a pause, and where can that traffic be stored? This guide provides a headroom worksheet, an incast calculation and a deployment test matrix for that decision.

The worked design is hypothetical: 32 servers, eight GPUs per server, and a selected 400 Gb/s Ethernet fabric. Calculations are executed arithmetic, not benchmark measurements or validated switch settings. The configuration examples are explicitly pinned to NVIDIA Cumulus Linux 5.9 documentation; they are not claims about the newest release or every Spectrum generation. Use your actual ASIC, network operating system and NIC support matrix before making changes.

1. The problem: three mechanisms, three different jobs

Remote Direct Memory Access over Converged Ethernet, or RoCE, needs an end-to-end design rather than a switch checkbox. The documented Cumulus lossless profile combines Priority Flow Control (PFC) and Explicit Congestion Notification (ECN), while its lossy profile uses ECN without enabling PFC for the RoCE priority.[3]

PFC is link-local flow control for a selected priority, not an application-aware congestion controller; it can pause a traffic class without pausing the entire Ethernet link.[3] ECN marks congestion for an endpoint response, and in RoCEv2 the receiver sends a Congestion Notification Packet (CNP) back toward the sender.[3] Buffer headroom is the space needed to absorb traffic that still arrives while the pause takes effect; NVIDIA's flow-control guidance explicitly considers interface delays and cable distance when calculating buffering.[1]

For this design, treat ECN as the primary congestion feedback mechanism and PFC as a safety mechanism for transient pressure. That is an engineering objective, not a promise that every workload will produce zero pause frames. A growing pause counter is a reason to correlate evidence, not automatic proof that the fabric is broken.

The practical buying question is therefore not simply “Does it support PFC?” Ask whether the supported profile, usable buffer partitioning, congestion telemetry and endpoint behavior meet your workload's completion-time objective under contention and failure.

RoCE data path with separate PFC and congestion feedback paths
Original diagram: Network freak / IPexpToBe. Generic conceptual illustration; no customer topology.

2. Define the 256-GPU example before sizing a buffer

Our hypothetical cluster contains 32 eight-GPU servers. For this worksheet, each server contributes one 400 Gb/s NIC port to the particular rail under review. Additional rails, if deployed, need their own port accounting and validation. We are not assuming that one NIC serves every GPU at full independent bandwidth, nor equating internal GPU interconnect bandwidth with Ethernet capacity.

The selected rail uses a routed leaf-spine fabric with sufficient nominal uplink capacity for the intended placement policy. That assumption does not prevent multiple sources from targeting one receiver simultaneously. We examine four senders converging on one 400 Gb/s destination path because it makes the difference between aggregate fabric capacity and receiver capacity explicit.

Keep these worksheet inputs with the design:

Input Hypothetical value Evidence required before deployment
Server population 32 servers, 256 GPUs Actual server and rail inventory
Selected rail link rate 400 Gb/s per server Negotiated link mode at both ends
Concurrent senders to one output Four Workload placement and traffic capture/telemetry
Effective pause reaction interval 1, 2 or 4 microseconds Platform guidance or validated timing model
Frame allowance used in arithmetic Two times 9,216 bytes Platform's actual frame and buffer-cell accounting
Number of modeled headroom reservations 32 Actual ingress priority-group allocation model
Lossless priorities One in this example Approved application-to-priority policy

These are design assumptions, not recommended defaults. A “9,216-byte allowance” here is an arithmetic input; it is not a Linux IP MTU command and does not specify how a particular ASIC counts Ethernet overhead. Do not silently substitute jumbo IP MTU, wire-frame length and cell-rounded buffer occupancy for one another.

3. Headroom worksheet: turn delay into bytes

Use the following first-order sensitivity model:

In-flight bytes = link rate in bits/second × effective reaction seconds ÷ 8

Worksheet allowance = in-flight bytes + explicit frame allowance

Define the effective reaction interval as the time from the local pause-trigger event until the last relevant incoming data reaches that buffer. It must include the applicable signaling, propagation, remote reaction and returning in-flight traffic effects. It is not just one-way cable propagation, and it is not automatically the end-to-end ECN feedback round trip.

For the table below, the frame allowance is two times 9,216 bytes. That is a deliberately visible illustrative margin, not an IEEE headroom formula, vendor-certified sizing tool or production minimum. ASIC pipelines, cell rounding, packet already in service, port mode, threshold semantics and shared headroom policies require platform-specific treatment. NVIDIA recommends leaving automatically derived flow-control thresholds and buffers alone unless the operator understands the lossless buffer requirements.[1]

Link rate 1 microsecond 2 microseconds 4 microseconds
100 Gb/s 30,932 B 43,432 B 68,432 B
200 Gb/s 43,432 B 68,432 B 118,432 B
400 Gb/s 68,432 B 118,432 B 218,432 B
800 Gb/s 118,432 B 218,432 B 418,432 B

All entries are calculated worksheet allowances using decimal bit rates and raw bytes. The 800 Gb/s row is a rate-sensitivity scenario, not an assertion that the documented software profile supports an arbitrary 800G switch.

At 400 Gb/s and four microseconds, the model contains 200,000 in-flight bytes plus 18,432 bytes of illustrative frame allowance. If a platform independently reserved 218,432 bytes for each of 32 modeled ingress priority groups, their arithmetic sum would be 6,989,824 bytes. That sum is not a switch-wide buffer requirement: it excludes other reservations and assumes an allocation model that might not match the hardware.

Reusable offline calculator

The following Python arithmetic was executed for this article. It does not interrogate a NIC or switch, and it does not validate a configuration. Replace inputs only after defining their units and provenance.

rates_gbps = (100, 200, 400, 800)
reaction_us = (1, 2, 4)
frame_allowance_bytes = 2 * 9216

for rate in rates_gbps:
    values = []
    for delay in reaction_us:
        in_flight = rate * 1_000_000_000 / 8 * delay / 1_000_000
        values.append(round(in_flight + frame_allowance_bytes))
    print(rate, values)

Use the table to challenge an assumption, not to overwrite a vendor profile. For example, if a proposed cable or port-mode change increases the validated reaction interval, rerun the sensitivity calculation and ask the vendor how that affects the supported headroom allocation. An unexplained “plenty of total buffer” answer is insufficient.

4. Why incast is not the same calculation

Now consider the four 400 Gb/s senders targeting a single 400 Gb/s output. In a deliberately simplified fluid model, total input is 1,600 Gb/s, output is 400 Gb/s, and net queue growth is 1,200 Gb/s while all sources sustain their offered rate.

Over an assumed ten-microsecond interval, that excess would add 1,500,000 bytes of backlog. This is theoretical offered-load arithmetic: it ignores packetization and assumes feedback has not reduced the sources. It does not describe a measured device or assert that all bytes occupy one physical buffer region.

This incast backlog is different from the per-ingress pause headroom calculation. Do not add both numbers mechanically to create a purchasing requirement, because allocation and accounting may overlap or reside in different resources. Use them to ask two separate questions: can the supported profile safely absorb the pause response, and how quickly does congestion feedback reduce persistent excess demand?

For a sustained workload, increasing buffer depth cannot turn one receiver into four. Consider spreading receivers, pacing application output, reducing concurrent senders, changing job placement or providing more destination capacity. Choose based on workload behavior, not a pause counter in isolation.

Four senders converging on one receiver in a hypothetical incast model
Original diagram: Network freak / IPexpToBe. Hypothetical 400 Gb/s incast arithmetic, not measured traffic.

5. Configuration workflow: establish classification before tuning

Cumulus maps packet markings into an internal switch priority used for classification and scheduling; that internal value is not itself written into the packet.[1] The documented RoCE profile maps switch priority 3 to RoCE traffic class 3 and priority 6 to CNP traffic class 6, with strict scheduling for the CNP class.[3] These are facts about the cited profile, not universal DSCP standards for every RoCE deployment.

Start with a signed-off mapping sheet covering sender marking, ingress trust, internal priority, lossless priority group, egress class and receiver handling. Add the reverse CNP path as a separate row. A configuration review that traces data but ignores feedback is incomplete.

Unexecuted lab example, Cumulus Linux 5.9 NVUE: the documentation provides explicit lossless mode configuration and operational inspection commands.[3]

nv set qos roce mode lossless
nv config apply
nv show qos roce
nv show interface swp16 qos roce status
nv show interface swp16 qos roce counters

Here, swp16 is an illustrative interface name, not an instruction to change a particular production port. Inspect the intended scope before applying a global profile. Capture configuration, reserve a maintenance window and ensure an out-of-band rollback path. NVIDIA warns that some QoS changes involving ASIC buffer modifications can cause momentary packet loss, even when using reload rather than restart.[1]

Do not copy generated RoCE configuration between ASIC families: the cited documentation explicitly says configuration generated for one Spectrum ASIC is not applicable to another.[3] Likewise, do not transplant an old file-based recipe into an NVUE-managed RoCE deployment; this release documents NVUE as the supported RoCE configuration path.[3]

My recommended workflow is: baseline one approved profile, validate classification, establish isolated performance, introduce controlled congestion, and only then consider tuning with the vendor. Change one variable per experiment and preserve the original profile as the comparison point.

6. Verification: counters must tell one coherent story

The documented interface counter view exposes RoCE ingress/egress activity, buffer usage and high-water marks, buffer discards, PFC pause packets and duration, ECN-marked packets and CNP activity.[3] Capture timestamped before-and-after snapshots rather than comparing unrelated lifetime totals.

For every trial, record offered load, receiver throughput, job duration, queue observations and endpoint congestion statistics together. Confirm the units and reset behavior of each counter before deriving rates. A quiet queue at the end of a burst does not invalidate an earlier high-water observation; retain both sampled time series and the platform's peak counters.

Use this proposed diagnostic matrix as an investigation guide, not a list of uniquely proven root causes:

Observation during a controlled test First question Next evidence to collect
RoCE bytes appear in the wrong class Was the marking trusted and mapped correctly? Sender settings plus ingress mapping on every hop
ECN marks rise but no expected CNP activity appears Is feedback generated and returned? Receiver NIC counters and reverse-path classification
CNP activity rises but load remains excessive Is the intended sender congestion control active? Sender firmware/driver settings and workload rate
Repeated long pause episodes Is the receiver or downstream path persistently overloaded? Receiver service rate, competing flows and queue timeline
Buffer discards coincide with pauses Is headroom, classification or allocation wrong? Port speed, cable model, pool configuration and priority-group peaks
Only cross-leaf trials degrade Is the shared fabric path the differentiator? Uplink distribution, class policy and link error deltas
Recovery follows a watchdog event Did protection sacrifice delivery to restore progress? Watchdog reason and documented recovery behavior

A CNP observation supports one part of the feedback chain; it does not by itself prove that the sender reacted correctly. Similarly, low pause activity is not a pass if application completion time is poor. Use the workflow-specific acceptance objective as the final arbiter.

7. Failure domains and alternative choices

Define failure containment before selecting thresholds. For this hypothetical cluster, treat a receiver, a leaf path, a rail and a shared congestion policy as separate domains. Losing a link can change the traffic distribution; a profile that worked only before the failure is not an accepted degraded-mode design.

Include an isolated test of pause propagation and the platform's supported PFC watchdog behavior. NVIDIA documents watchdog support in its QoS guide.[1] Ask exactly what traffic is dropped, which priority is affected and how recovery occurs on your platform; do not describe watchdog activation as a lossless success. Never deliberately generate a pause storm on a production shared fabric.

There are three useful design alternatives:

  • Supported lossless RoCE profile: choose this when the NIC/switch combination and operational team can validate classification, PFC containment and congestion feedback together.
  • Supported lossy RoCE profile: Cumulus documents an ECN-based lossy mode without enabling PFC for the RoCE priority.[3] Evaluate actual workload recovery, tail latency and completion time before selecting it; “no PFC” is not a free performance improvement.
  • Dedicated fabric or another interconnect: consider separating traffic or evaluating InfiniBand when operational isolation is more important than sharing Ethernet infrastructure. Compare total operational cost and measured job behavior, not unrelated bandwidth headline numbers.

Keep backup bursts and management accessibility in the test plan. Do not automatically assign unrelated traffic to the lossless class. Our recommendation is to admit traffic deliberately and verify that the management path remains usable during the worst approved congestion experiment.

8. Deployment and acceptance checklist

The following is a proposed acceptance contract. Agree exact throughput and recovery tolerances with the workload owner before testing; there is no invented benchmark target here.

Test Scope Evidence to retain Acceptance decision
Inventory and compatibility Every NIC and switch family Versions, port modes, supported profile No unsupported combinations or unexplained drift
Mapping audit Data and CNP paths Marking-to-priority-to-class worksheet Intended mapping on every tested hop
Isolated baseline Representative server pairs Throughput, latency, error deltas Meets agreed hardware/workload baseline
Controlled incast Same-leaf and cross-leaf Queue peaks, marks, pauses, CNPs, job time Meets agreed job-time bound with explained counters
Mixed workload Compute plus storage/background traffic Per-class behavior and management reachability Isolation objectives remain satisfied
Single-link failure Each relevant path type Recovery time and post-failure distribution Meets degraded-mode target without persistent stalls
Feedback fault investigation Isolated lab only Endpoint and reverse-path observations Failure is detectable and rollback works
Configuration rollback Approved maintenance test Before/after state and restored baseline Known-good policy and workload behavior restored

Run repeated trials with identical placement and payload settings before attributing a difference to a buffer change. Preserve failures as well as successes. If only one favorable run is retained, the evidence package cannot establish repeatability.

A useful handover folder contains the mapping sheet, platform-specific headroom explanation, executed worksheet inputs, test topology, software versions, counter snapshots, workload results and rollback procedure. That makes the asset reusable when a cable type, firmware version or job mix changes.

Practical takeaways

Size the reaction window, not just the switch's advertised memory. The worksheet makes rate and delay assumptions visible, while the incast example explains why buffer depth alone cannot solve sustained oversubscription. Use a supported ASIC-specific profile first, validate the endpoint feedback loop, and approve changes against job completion and failure recovery—not a single counter.

For adjacent work, use the AI GPU Cluster Network Calculator for port budgets, the NCCL Slow AllReduce test matrix for workload localization, and the 256-GPU acceptance checklist for handover. The AI Infrastructure hub and Data Center hub collect the wider series.

Sources

Comments

0 Responses to "RoCE PFC Buffer Sizing Guide: 400G Headroom Worksheet and ECN Checklist"

Post a Comment

Popular Posts