Paper Detail

VibeVoice-ASR-Streaming Technical Report

Yujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu, Yaoyao Chang, Wenhui Wang, Zilong Wang, Zehua Wang, Yan Xia, Jiajun Zhang, Xie Chen, Furu Wei

arxiv Score 6.3

Published 2026-09-02 · First seen 2026-09-03

General AI

Abstract

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{tu2026vibevoice,
  title = {VibeVoice-ASR-Streaming Technical Report},
  author = {Yujie Tu and Zhiliang Peng and Jianwei Yu and Li Dong and Songchen Xu and Yaoyao Chang and Wenhui Wang and Zilong Wang and Zehua Wang and Yan Xia and Jiajun Zhang and Xie Chen and Furu Wei},
  year = {2026},
  abstract = {Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming spea},
  url = {https://arxiv.org/abs/2609.02812},
  keywords = {eess.AS, speaker-attributed ASR, end-to-end, LLM-based, streaming, speaker diarization, lookahead audio, WER/CER, huggingface daily},
  eprint = {2609.02812},
  archiveprefix = {arXiv},
}

Metadata

{}