Nvidia's latest solution for soaring enterprise costs: NeMo Switchyard software router

https://image.theregister.com/5286938.jpg?imageId=5286938&x=0&y=0&cropw=100&croph=100&panox=0&panoy=0&panow=100&panoh=100&width=1200&height=683

ai and ml

NeMo Switchyard brings GPT-5-style model routing to the mainstream

Soaring AI infrastructure costs and model pricing, combined with uncertain returns on investment, threaten to stall enterprise adoption.

To make enterprise AI spend a bit more manageable, Nvidia this week unveiled a new software platform that blurs the line between expensive proprietary models and open weights alternatives.

Announced alongside Nemotron 3.5-30B-A3B-Lightning, Nvidia’s latest open weights model, NeMo Switchyard is the GPU giant’s latest overture to enterprise. So what exactly is it? Well, it’s a router.

The idea is simple. Switchyard essentially functions as a proxy that sits between the inference server’s API endpoint and the models. But rather than sending every request to the same model, Switchyard can be configured to route prompts to different models in order to optimize for cost, latency, or output quality.

By routing some requests to smaller, cheaper, and potentially locally hosted...

Copyright of this story solely belongs to theregister.com. To see the full text click HERE