Paper Detail
Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou
Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), WanSong is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.
No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.
No ranking explanation is available yet.
No tags.
@misc{chen2026wansong,
title = {WanSong v1.0 Technical Report},
author = {Binghui Chen and Pandeng Li and Yu Liu and Jingren Zhou},
year = {2026},
abstract = {Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\textbackslash{}eg, AR followed by diffusion), WanSong is a pure diffusion-based model that directly generate},
url = {https://huggingface.co/papers/2607.14749},
keywords = {huggingface daily},
eprint = {2607.14749},
archiveprefix = {arXiv},
}
{}