Skip to main content
  1. Index/

RoCE (RDMA over Converged Ethernet)

RoCE (RDMA over Converged Ethernet) implements RDMA semantics on Ethernet (RoCEv2 uses UDP/IP), so NICs can perform remote memory access with low CPU utilization over the same physical switches many enterprises already operate. Its objective is to deliver InfiniBand-like GPU communication economics—fast NCCL all-reduces, NVMe-oF, llm-d KV moves—without maintaining a separate InfiniBand fabric. RoCE requires lossless Ethernet behavior: Priority Flow Control (PFC), Explicit Congestion Notification (ECN), buffer tuning, and often dedicated traffic classes so RDMA traffic is not dropped under burst load.

Architecturally, RoCEv2 is routable L3 Ethernet; InfiniBand uses different link and subnet management. Misconfigured switches cause silent performance cliffs (retransmits, collapse to slow paths), so RoCE is as much a datacenter design choice as a NIC feature. Within a server, NVLink still handles GPU peer traffic; RoCE links nodes. The CPU role matches other RDMA: setup and orchestration, not per-byte copying. AI workloads do not distinguish RoCE vs IB at the PyTorch API—NCCL selects transports based on environment and hardware—but ops teams must qualify firmware, cable, and QoS end to end.

Red Hat documents RoCE on RHEL (ConnectX drivers, rdma-core, tuning guides) and OpenShift networking for accelerated workloads, overlapping telco core and AI cluster designs. Customers who standardize on Ethernet for cost and operations use RoCE for GPU scale-out while running OpenShift AI or bare-metal training clusters on supported RHEL. Red Hat support focuses on kernel, driver, and platform integration; switch QoS profiles remain a joint validation exercise with network vendors.

Related

InfiniBand

InfiniBand is a high-performance network fabric designed for datacenter and HPC clusters, natively supporting RDMA (Remote Direct Memory Access) with low latency, high bandwidth, and features such as adaptive routing and congestion control at the link layer. Its objective in AI is to connect many GPU servers so distributed training (gradient all-reduce via NCCL) and multi-node inference (tensor parallel, llm-d prefill/decode KV transfer) are not limited by TCP overhead on a CPU. InfiniBand NICs (e.g. NVIDIA ConnectX) present verbs APIs; subnets are managed with an Subnet Manager and partitioned for multi-tenant isolation.

RDMA (Remote Direct Memory Access)

RDMA (Remote Direct Memory Access) allows a network adapter to transfer data between the memory of two machines with little CPU overhead, low latency, and often kernel bypass (userspace stacks such as verbs on InfiniBand or RoCE). Its objective in AI infrastructure is to keep GPUs fed and synchronized: distributed training exchanges gradients quickly, disaggregated inference (llm-d) moves KV cache blocks between prefill and decode nodes, and NVMe-oF storage delivers checkpoints without the host spending cycles copying every byte. DPUs and SmartNICs also use RDMA paths for storage and east-west traffic while the host CPU runs models.

NCCL (NVIDIA Collective Communications Library)

NCCL (NVIDIA Collective Communications Library) implements collective operations—all-reduce, broadcast, reduce-scatter, all-gather, and others—optimized for NVIDIA GPUs across NVLink within a node and RDMA (InfiniBand or RoCE) across nodes. Its objective in AI is to make distributed training and multi-GPU inference (tensor parallelism) scale: gradient shards must merge every step; attention and MLP partitions must exchange activations with minimal latency. Frameworks (PyTorch DDP/FSDP, vLLM tensor parallel) call NCCL (or delegate to it) rather than hand-rolling socket code.