Control How Your GPU Shares Work with Green Contexts
Oct 6, 2026, 8:00 AM · NVIDIA Developer

NVIDIA's green contexts let one program carve a GPU into dedicated lanes, and the latency win for urgent work is big enough that inference and robotics teams should take notice.
Why it matters
Modern GPU apps rarely run one job. A latency-sensitive kernel often shares the chip with heavy background work, and today's tools don't give developers much control over who gets which compute units. NVIDIA's answer is green contexts: a way for an application to reserve a subset of a GPU's streaming multiprocessors and workqueues within a single process.
The feature has existed in the Driver API since CUDA 12.4. What's new is that CUDA 13.1 exposes it through the Runtime API, the layer most developers actually use. In NVIDIA's own test on a 148-SM Blackwell GPU, a small critical kernel ran in 0.007 milliseconds with eight SMs set aside for it, versus 0.140 ms with stream priority alone and 3.727 ms with no priority at all.
From the desk
This is plumbing, and we think plumbing like this matters more than it gets credit for. A lot of real-world AI isn't one giant model saturating a chip. It's a pipeline: preprocessing next to inference, communication overlapping with matrix math, sensor processing that has to react now. When those pieces step on each other, the result is jitter, and jitter is what makes a system feel unreliable.
The key insight in NVIDIA's write-up is that stream priority has a ceiling. A high-priority kernel still has to wait for already-running blocks to drain off an SM, because priority can't preempt them. Green contexts skip the wait by keeping some SMs off-limits to the bulk job. That's the whole trick, and the numbers show it's a strong one.
We'd keep the tradeoff in view. Reserving SMs for urgent work means fewer SMs for throughput, so teams are trading peak utilization for predictability. Those numbers come from a vendor-designed benchmark built to show the effect. Real workloads will land somewhere less dramatic, and getting partitions right will take tuning.
We like that it's opt-in and additive. Existing code keeps targeting the full device, and teams can adopt partitions one component at a time. If this becomes standard practice, the payoff is AI systems with tighter, more predictable response times on shared hardware — and a little more lock-in to NVIDIA's programming model along the way.
Context
CUDA contexts were designed when GPUs were smaller and usually ran one dominant workload. NVIDIA describes them as heavyweight, with context-switch overhead; green contexts are lightweight to create and destroy, and doing so doesn't synchronize unrelated GPU work.
Who feels it
- Inference and training engineers
- A practical tool for overlapping communication with compute, or protecting latency-critical operators from background jobs.
- Robotics and edge teams
- Sensor-processing pipelines that need kernels to start immediately get a more deterministic option.
- Framework maintainers
- Runtime API support makes it easier to build partition-aware scheduling into higher-level libraries.
What to watch
- Adoption of green contexts in major inference and training frameworks
- Independent measurements on production workloads rather than synthetic tests
- Whether AMD and other accelerator vendors expose similar partitioning controls