Paper Detail

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Yan Zhan, Mengkai Hou, Wanting Zhang, Zhijun Gao

huggingface Score 6.0

Published 2026-08-23 · First seen 2026-08-26

General AI

Abstract

Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under-explored despite its direct effect on coverage and redundancy. We introduce Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill. Relative to a baseline using training corpus length statistics, EV recovers 64.9%, 65.3%, and 33.0% of the COMET-22 gain from reference target lengths on EntoZh, ZhtoEn, and EntoDe. Our diagnostics show that denoising-friendly lengths need not match reference lengths. Evaluation by three translation experts supports the EnleftrightarrowZh adequacy gains, with stronger evidence on ZhtoEn. Compared with a LLaMA-3-8B autoregressive (AR) model trained on the same fine-tuning data, the EV system ties on EntoZh and leads on ZhtoEn; an oracle-length diagnostic further shows that, in this masked diffusion MT setting, deciding which tokens to reveal first matters less than how the target length is supplied.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{zhan2026length,
  title = {Length-Adaptive Decoding for Masked Diffusion Machine Translation},
  author = {Yan Zhan and Mengkai Hou and Wanting Zhang and Zhijun Gao},
  year = {2026},
  abstract = {Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under-explored despite its direct effect on coverage and redundancy. We introduce Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entrop},
  url = {https://huggingface.co/papers/2608.22274},
  keywords = {masked diffusion language models, dLLMs, entropy-valley, predictive entropy, all-mask forward passes, COMET-22, denoising-friendly lengths, autoregressive, masked diffusion MT, code available, huggingface daily},
  eprint = {2608.22274},
  archiveprefix = {arXiv},
}

Metadata

{}