NVIDIA NemotronLabs VoiceChat 11B: Open, Real-time S2S AI
Summary
NVIDIA has released NemotronLabs VoiceChat 11B, an open speech-to-speech model designed for real-time, full-duplex conversations. This new model integrates speech understanding and generation into a single network, avoiding the need to chain multiple systems. It achieves a smooth turn-taking latency of 448 milliseconds. The model can also listen while speaking, allowing users to interrupt, and the agent will yield with a 1.00 take-over rate at 480 milliseconds. Notably, it's the first open full-duplex model to support tool calling while maintaining conversational flow. While deployable for pilots, NVIDIA states it's currently for research purposes only. There are known issues, including a two-minute audio context limit and potential degradation. Companies with a GPU of at least 80 GB of VRAM can use it for applications like contact centers, in-cabin assistants, and live lookup. This technology could significantly improve the naturalness and efficiency of voice-based interactions.
This is an AI-generated audio summary. Always check the original source for complete reporting.