How DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software Stack | NVIDIA Technical Blog

How DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software Stack | NVIDIA Technical Blog

By Elizabeth Goodman
Publication Date: 2026-10-06 19:07:00

GPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services. When the CPU sits in the middle of every network transaction, it becomes a bottleneck on the critical path, adding latency and limiting how efficiently distributed applications can respond in real time.

NVIDIA DOCA GPUNetIO, a GPU-centric networking SDK layer for real-time packet processing and data movement, addresses this directly. It brings together technologies such as GPUDirect RDMA, GPUDirect Async Kernel-Initiated (GDA-KI) and GDRCopy so CUDA kernels can directly drive Ethernet, RDMA, Verbs, and DMA operations while keeping the CPU out of the application critical path.

Specifically, DOCA GPUNetIO provides both:

  • CPU functions on the control path to export on GPU memory transport objects created with DOCA Ethernet, DOCA Verbs, DOCA DMA, DOCA CommChannel
  • GPU CUDA functions on the data path to allow the creation of CUDA kernels that can manipulate the transport objects exported

Since NVIDIA first introduced GPU-centric packet processing with DOCA GPUNetIO, the framework has matured significantly, evolving from a powerful way to remove the CPU from the critical path into a unified GDA-KI foundation integrated across NVSHMEM, NCCL, Aerial 5G SDK, UCX/NIXL, NVQLink with Holoscan Sensor Bridge operator , Holoscan Advanced Network Operator , DeepEP/HybridEP, and others.

This post covers what’s new: the…