NVIDIA VoiceChat-11B Brings Full-Duplex AI Speech to Real-Time Agents
Overview
NVIDIA-NemotronLabs-VoiceChat-11B is an 11-billion parameter end-to-end full duplex speech model developed by nvidia that performs streaming speech understanding and generation in a single unified architecture. Unlike traditional cascaded pipelines that chain ASR, LLM, and TTS models separately, this model operates directly on audio signals with 16 kHz input and produces 22.05 kHz output, achieving ~450 milliseconds turn-taking latency while supporting real-time interruption handling and tool calling. The architecture combines a Fast Conformer speech encoder, a Nemotron Nano V2 9B LLM backbone, and a proprietary NVIDIA TTS decoder with a separate output channel for tool-calling scripts. Training involved approximately 550,000 hours of audio across a hybrid blend of real speech datasets (Fisher, LibriVox, LibriTTS, VCTK) and synthetic data generated from text corpora including Nemotron 5.5, Ultrachat, and PromptTTS. The model runs on NVIDIA GPU-accelerated systems via the vLLM inference engine, supporting A100, H100, H200, B100, B200, and RTX-6000 hardware on...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE