Paper Detail

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen, Zhehuai Chen, Zijia Chen, Francesco Ciannella, Slyne Deng, Mikyas Desta, Harishchandra Dubey, Slim Essid, Nourchene Ferchichi, Boris Ginsburg, Mariana Graterol Fuenmayor, Negar Habibi, Kevin Hu, Anand Joseph, Viraj Karandikar, Myungjong Kim, Viacheslav Klimkov, Seelan Lakshmi Narasimhan, Lily Lee, Jason Li, Eileen Long, Ameya Mahabaleshwarkar, Aditya Malte, Adi Margolin, Sasha Meister, Valentin Mendelev, Oluwatobi Olabiyi, Ankita Pasad, Yifan Peng, Elena Rastorgueva, Jayda Ritchie, Jason Roche, Nikhil Srihari, Yuanhang Su, Yoshi Suhara, Viet Anh Trinh, Jinhan Wang, Piotr Zelasko, Hui Wang, Puhui Meng, Chaosen Zhang, Yunsheng Liu, Shawn Wang, Wenjing Li, Zhonglei He

arxiv Score 12.8

Published 2026-09-18 · First seen 2026-09-21

General AI

Abstract

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100\% takeover following user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes its response after user backchannels in 93\% of cases. NemotronLabs VoiceChat obtains a 55.1 normalized average on VoiceBench and, on Full-Duplex-Bench 3.0 (FDB 3.0), achieves 82.5\% tool-selection F1, while argument accuracy and end-to-end tool execution remain areas for improvement. These results demonstrate that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{balam2026nemotronlabs,
  title = {NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities},
  author = {Jagadeesh Balam and Travis Bartley and Edresson Casanova and Sanjay Chauhan and Chen Chen and Zhehuai Chen and Zijia Chen and Francesco Ciannella and Slyne Deng and Mikyas Desta and Harishchandra Dubey and Slim Essid and Nourchene Ferchichi and Boris Ginsburg and Mariana Graterol Fuenmayor and Negar Habibi and Kevin Hu and Anand Joseph and Viraj Karandikar and Myungjong Kim and Viacheslav Klimkov and Seelan Lakshmi Narasimhan and Lily Lee and Jason Li and Eileen Long and Ameya Mahabaleshwarkar and Aditya Malte and Adi Margolin and Sasha Meister and Valentin Mendelev and Oluwatobi Olabiyi and Ankita Pasad and Yifan Peng and Elena Rastorgueva and Jayda Ritchie and Jason Roche and Nikhil Srihari and Yuanhang Su and Yoshi Suhara and Viet Anh Trinh and Jinhan Wang and Piotr Zelasko and Hui Wang and Puhui Meng and Chaosen Zhang and Yunsheng Liu and Shawn Wang and Wenjing Li and Zhonglei He},
  year = {2026},
  abstract = {We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming },
  url = {https://arxiv.org/abs/2609.21967},
  keywords = {cs.CL, cs.AI},
  eprint = {2609.21967},
  archiveprefix = {arXiv},
}

Metadata

{}