Nvidia's Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests

https://images.ctfassets.net/jdtwqhzvc2n1/6qbSNxUnm1pPACzx2jzAAQ/18b1af901de9038b19d1069310092946/lightning-nemotron-smk1.jpg?w=800&q=75

Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its own engineering project, one that has to be maintained every time a workflow changes.

Nvidia is proposing a fix that touches both ends of that problem at once. The company is out on Tuesday with Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model built for high-volume, specialized agent tasks, alongside NeMo Switchyard, an open-source library that routes each step of an agent workflow to whichever model fits it best.

The headline numbers: According to Nvidia, Lightning delivers up to 4x faster output than comparable models in its class, completing agentic tasks roughly 30% faster than Qwen3.6-35B at matching accuracy. Paired through Switchyard, Nvidia says the combination holds frontier-level task completion while cutting benchmark...

Copyright of this story solely belongs to venturebeat.com. To see the full text click HERE