By Elizabeth Goodman
Publication Date: 2026-10-06 15:00:00
GPU applications increasingly consist of multiple independent components running at the same time within a single process: a latency-sensitive operator alongside a throughput-oriented background kernel; a data preprocessing stage alongside model inference; or multiple stages of a processing workflow sharing a single GPU.
Controlling how GPU resources are shared between them remains difficult. Components can interfere unpredictably, and existing tools offer limited ability to partition resources.
Green contexts address this by letting applications explicitly select a subset of GPU execution resources and target work to those resources directly. Traditional CUDA contexts weren’t designed for this usage model. They are heavyweight, incur hardware context-switch overhead, and reflect assumptions from an earlier era, when GPUs were smaller and applications typically ran as a single dominant workload.
Green contexts have been available in the Driver API since NVIDIA CUDA 12.4. Starting with CUDA 13.1, they are also accessible through the Runtime API, allowing an application, within its process, to explicitly define where work runs and how execution resources are divided.
How green contexts work
One primary use of green contexts is SM partitioning. By assigning a specific subset of SMs to a green context, applications can target work submitted through that green context to those SMs. This can enable multiple workloads to run concurrently on the GPU without…

