By Tobias Mann
Publication Date: 2026-08-12 19:00:00
ai and ml
NeMo Switchyard brings GPT-5-style model routing to the mainstream
Soaring AI infrastructure costs and model pricing, combined with uncertain returns on investment, threaten to stall enterprise adoption.
To make enterprise AI spend a bit more manageable, Nvidia this week unveiled a new software platform that blurs the line between expensive proprietary models and open weights alternatives.
Announced alongside Nemotron 3.5-30B-A3B-Lightning, Nvidia’s latest open weights model, NeMo Switchyard is the GPU giant’s latest overture to enterprise. So what exactly is it? Well, it’s a router.
The idea is simple. Switchyard essentially functions as a proxy that sits between the inference server’s API endpoint and the models. But rather than sending every request to the same model, Switchyard can be configured to route prompts to different models in order to optimize for cost, latency, or output quality.
By routing some requests to smaller, cheaper, and potentially locally hosted AI models, Nvidia claims Switchyard can cut job completion costs by 74 percent relative to using Claude Opus 4.8 alone, albeit with an approximately six-point accuracy tradeoff.
The right tool for the job
The key metric in all of this is completion cost rather than price per token. A model might cost one-tenth as much as OpenAI’s or Anthropic’s top model, but if it requires 10x the tokens to complete the…



