InfiniBand vs RoCEv2 for Inference: Start With the Traffic, Not the Logo
Choose an inference fabric from traffic shape, tail latency, failure recovery, and operator burden—not peak bandwidth or a switch logo.
My current inference machine is one Threadripper PRO 7960X with one RTX PRO 6000. It is not a cluster, and I have not run an InfiniBand-versus-RoCE bake-off. That limitation exposes the first mistake in many fabric debates: arguing about the scale-out network before proving that the workload must leave the box.
I come to this from designing SASE and RoCEv2 networks, with experience around InfiniBand-adjacent infrastructure. The lesson that carries over is unglamorous. A protocol name does not rescue a poor traffic model. Map the flows first. Then decide what deserves measurement.
The practical thesis is simple: keep communication inside the fastest scale-up domain available, and choose a scale-out fabric only after the workload, latency objective, congestion pattern, failure model, and operator capacity are explicit. InfiniBand and RoCEv2 can both carry serious AI workloads. They impose different operating costs, and neither fixes an architecture that crosses nodes unnecessarily.
Find the Real Network Boundary
A serving system may contain several networks, even when a diagram labels all of them “the fabric.” A single GPU does not perform a distributed collective. Multiple GPUs inside one server may communicate over PCIe or NVLink. Only traffic that crosses a host reaches the InfiniBand or Ethernet scale-out network.
That distinction matters because decode can be sensitive to communication that never touches an external switch. NVIDIA tested Gemma 2 inference with eight-way tensor parallelism on one node containing eight H100 GPUs connected by NVLink. An all-reduce of roughly 30 KB accounted for about 23% of end-to-end decode latency because dependencies prevented overlap with nearby compute. A custom one-shot collective, fused with adjacent operations, improved end-to-end decode latency by about 27%.[4]
The result demonstrates the cost of small, synchronous communication. It does not demonstrate an InfiniBand or RoCEv2 advantage. NVIDIA specifies 900 GB/s of GPU-to-GPU bandwidth for the eight-GPU DGX H100/H200 NVSwitch domain used for this class of scale-up communication.[5] Buying a 400 Gb/s external fabric would not improve a request path that remains inside that domain.
Start the design review with a boundary map:
- Which parallelism mode creates each flow: tensor, pipeline, expert, data, or request routing?
- Does the flow stay within a GPU, cross GPUs inside a host, or cross hosts?
- What is the message-size distribution, not merely its average?
- Which transfers sit on the token-generation critical path?
- Which flows can overlap with compute, and which force every rank to wait?
Cross a node boundary because model capacity, throughput, fault-domain design, expert parallelism, or prefill/decode separation requires it. Do not cross it because the network purchase came first.
Inference Is More Than One Packet Profile
Training guidance often assumes large synchronized collectives where sustained bandwidth dominates. Inference produces a less uniform mix. Tensor-parallel decode can generate frequent small collectives with synchronization in the critical path. Expert parallelism can create all-to-all traffic with uneven destinations. Disaggregated serving moves KV cache between prefill and decode workers.
vLLM’s disaggregated-prefill implementation illustrates the last case. It runs separate prefill and decode instances, then uses a connector to transfer KV cache and results between them. The feature remains experimental, and its documentation explicitly states that disaggregated prefill does not improve throughput by itself.[8] Its value is isolation: teams can tune time to first token and inter-token latency separately, prevent long prefills from blocking decode, or use different parallel strategies for each stage.
That design also adds another network dependency. KV transfer has a different size and burst pattern from a small all-reduce. It may create incast at decode workers, compete with collectives, or turn a slow connector into the request bottleneck. A network selected from one bandwidth chart cannot represent all of these paths.
For each traffic class, record four things before testing a fabric: payload distribution, concurrency, synchronization behavior, and acceptable tail latency. Add background traffic and placement constraints. A clean-room benchmark with one flow can validate wiring, but it cannot predict a shared serving cluster.
What the Large Clusters Actually Prove
Meta built two clusters with 24,576 H100 GPUs each. One used a RoCE fabric and the other NVIDIA Quantum-2 InfiniBand; both connected 400 Gb/s endpoints.[1] Meta later reported that, after tuning, the two clusters delivered equivalent performance for its large generative-AI training workloads, and that it used the RoCE cluster to train the largest Llama 3 model.[2]
This is strong evidence against the claim that Ethernet cannot support large AI systems. It is not evidence that default RoCE settings equal a tuned InfiniBand deployment, nor is it an inference benchmark.
The tuning history carries the useful lesson. Meta reported initial communication utilization ranging from 10% to 90% on unoptimized large clusters, compared with more than 90% after changes across scheduling, routing, collective software, and the network.[1] In its RoCE engineering report, fragmented job placement and flow collisions degraded training performance by more than 30% in one case. Enhanced ECMP combined with queue-pair scaling improved AllReduce performance by up to 40% over baseline ECMP in that environment.[3]
Meta also found that default DCQCN settings performed poorly in its 400G deployment. Firmware changes created bugs and reduced visibility into congestion-notification counting, so the team proceeded without DCQCN and coordinated traffic admission through the collective library and receiver behavior instead.[3] That is not a recipe to copy blindly. It shows that congestion control, collective implementation, placement, and routing form one system.
The narrow conclusion is operational: RoCE performance is produced, not inherited. A successful deployment depends on path entropy, queue design, PFC and ECN policy, buffers, NIC firmware, collective behavior, topology-aware scheduling, and telemetry. A team that owns those layers may value Ethernet’s flexibility. A team that does not must price the engineering gap.
Failure Surfaces Differ, but Neither Fabric Is Automatic
RoCEv2 carries RDMA over routed Ethernet. That allows reuse of familiar switching, cabling, automation, and routing practices. It also exposes Ethernet-specific failure modes that can preserve link-up status while damaging latency.
NVIDIA’s Cumulus Linux documentation describes RoCE configuration as complex and documents lossless operation using PFC and ECN. The same interface exposes pause packets, pause duration, ECN-marked packets, queue use, and discards because those signals are part of operating the service, not optional switch trivia.[6] PFC can protect a lossless traffic class, but a bad threshold or blocked receiver can spread pauses. ECN may mark congestion without preventing a queueing tail if endpoints react too slowly or inconsistently.
NVIDIA’s NCCL troubleshooting guide separates bandwidth from latency for the same reason: bandwidth can look healthy while tail latency remains poor. It recommends point-to-point latency tests and inspection of PFC, ECN, CNP, queue drops, retransmission symptoms, and RDMA errors when stalls or long tails appear.[7] An operator needs correlated timestamps from the request layer, collective library, NIC, and switches to determine whether a token stall came from compute, synchronization, congestion, loss recovery, or a degraded link.
InfiniBand narrows the Ethernet tuning surface and offers an integrated RDMA fabric. It does not eliminate operations. The subnet manager, routing, cables, rails, firmware, PCIe topology, GPU Direct configuration, and collective settings remain possible failure points. NVIDIA’s guidance calls for checking the subnet manager, link state and rate, port-error counters, topology, and routing, then comparing bandwidth and latency rather than trusting one test.[7]
The real comparison is therefore not “complex versus automatic.” It is one failure surface versus another, measured against the skills, tooling, vendor support, and change process already present in the organization.
Run a Serving Bake-Off, Not a Link Contest
A useful proof of concept holds the serving workload constant and changes only the fabric-specific stack. Match GPU model and count, host topology, NIC placement, transceivers, switch oversubscription, model, quantization, parallelism, software versions, request trace, and background load. Record every deviation.
Begin with component tests to catch basic faults:
- Verify link state, negotiated rate, PCIe width, GPU Direct behavior, and rail mapping.
- Measure point-to-point bandwidth and latency across every path class, including cross-rack paths.
- Run collective tests across the real message-size distribution, not only large buffers that approach line rate.
- Inject contention, a failed link, a degraded link, and an unavailable switch path.
Then replay the serving workload. Report time to first token, inter-token latency, end-to-end request latency, and goodput at the target concurrency. Show P50, P95, and P99 rather than one average. Segment results by prompt and output length so changes in traffic mix do not masquerade as fabric effects.
Correlate application results with collective duration by message size, NIC throughput, queue depth, PFC pause duration, ECN marks, CNPs, drops, retransmission or retry counters, and GPU idle time. Also measure recovery: requests lost during a fault, time to restore capacity, and performance after rerouting. A fabric that wins an empty-network throughput run but produces unstable P99 under contention has not won the serving decision.
Set acceptance criteria before the test. For example, require both candidates to stay within the latency service-level objective at expected concurrency, recover from a single-link failure without manual repair, and expose enough telemetry to identify the cause of a tail regression. Compare operator hours and unresolved incidents alongside hardware cost. Otherwise the cheaper bill of materials can become the more expensive service.
A Conditional Decision Rule
For a single node, buy no scale-out fabric for inference. Measure the PCIe or NVLink path and optimize the collective implementation first.
For latency-sensitive tensor parallelism that must cross nodes, give InfiniBand the first proof of concept when the team wants a more integrated fabric and lacks deep lossless-Ethernet operating experience. Make RoCEv2 prove equivalent P99 inter-token latency under contention, not merely similar peak bandwidth.
For throughput-oriented serving, disaggregated prefill/decode, or an organization with strong Ethernet automation and telemetry, give RoCEv2 a controlled trial. Meta’s published work establishes that the performance ceiling can be high. The same work documents how much full-stack engineering may be required to reach it.[1][2][3]
Ultra Ethernet should influence the next test plan, not overwrite current evidence. The UEC 1.0 specification defines transport profiles with multiple ordering modes, congestion-management mechanisms, multipath path selection, and packet trimming.[9] Those mechanisms target real limitations in conventional Ethernet transports, but a specification is not an implementation benchmark. Validate the shipping NIC, switch, software, and observability stack against the workload before assigning production value.
The purchase decision follows the request path. Keep communication within the scale-up domain when the model permits. When traffic must cross nodes, evaluate the actual messages, contention, tails, failure recovery, and operator burden. The logo on the switch is an input to that work, not the answer.
Sources
[1] https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure
[2] https://engineering.fb.com/2024/06/12/data-infrastructure/training-large-language-models-at-scale-meta
[3] https://engineering.fb.com/2024/08/05/data-center-engineering/roce-network-distributed-ai-training-at-scale
[4] https://developer.nvidia.com/blog/optimizing-for-low-latency-communication-in-inference-workloads-with-jax-and-xla
[5] https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-to-dgxh100.html
[6] https://docs.nvidia.com/networking-ethernet-software/cumulus-linux-510/Layer-1-and-Switch-Ports/Quality-of-Service/RDMA-over-Converged-Ethernet-RoCE
[7] https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/troubleshooting/networking_troubleshooting.html
[8] https://docs.vllm.ai/en/stable/features/disagg_prefill
[9] https://ultraethernet.org/wp-content/uploads/sites/20/2025/06/UE-Specification-6.11.25.pdf
Share this article
Comments
Comments are powered by GitHub Discussions via Giscus.
To enable comments, configure Giscus at giscus.app and update the Comments component with your repo settings.