By Elizabeth Goodman
Publication Date: 2026-10-06 19:07:00
GPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services. When the CPU sits in the middle of every network transaction, it becomes a bottleneck on the critical path, adding latency and limiting how efficiently distributed applications can respond in real time.
NVIDIA DOCA GPUNetIO, a GPU-centric networking SDK layer for real-time packet processing and data movement, addresses this directly. It brings together technologies such as GPUDirect RDMA, GPUDirect Async Kernel-Initiated (GDA-KI) and GDRCopy so CUDA kernels can directly drive Ethernet, RDMA, Verbs, and DMA operations while keeping the CPU out of the application critical path.
Specifically, DOCA GPUNetIO provides both:
- CPU functions on the control path to export on GPU memory transport objects created with DOCA Ethernet, DOCA Verbs, DOCA DMA, DOCA CommChannel
- GPU CUDA functions on the data path to allow the creation of CUDA kernels that can manipulate the transport objects exported
Since NVIDIA first introduced GPU-centric packet processing with DOCA GPUNetIO, the framework has matured significantly, evolving from a powerful way to remove the CPU from the critical path into a unified GDA-KI foundation integrated across NVSHMEM, NCCL, Aerial 5G SDK, UCX/NIXL, NVQLink with Holoscan Sensor Bridge operator , Holoscan Advanced Network Operator , DeepEP/HybridEP, and others.
This post covers what’s new: the…

/Nvidia%20logo%20and%20sign%20on%20headquarters%20by%20Michael%20Vi%20via%20Shutterstock.jpg?resize=1600,1067&ssl=1)
