Paper Detail

From Global to Factor-Wise Expert Composition in Discrete Diffusion Models

Haozhe Huang, Yudong Xu, Abhijoy Mandal, Alán Aspuru-Guzik

arxiv Score 12.2

Published 2026-07-13 · First seen 2026-07-14

General AI

Abstract

Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods are fundamentally limited by working on a per-sample basis, treating each generated state monolithically and ignoring the potential spatial or functional specializations of different experts. In this work, we address this limitation by proposing FactorDiff - a factor-wise composition framework for diffusion models. We posit that samples can be further decomposed into smaller factors, and propose a sampling process that dynamically routes each factor to the most relevant expert. We instantiate this framework with spatial/pixel-level compositions and validate it on the ARC-AGI benchmark, demonstrating that simple factor-specific routing consistently outperforms complex global scalar weighting schemes on tasks that require logical consistency and spatial disentanglement.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{huang2026global,
  title = {From Global to Factor-Wise Expert Composition in Discrete Diffusion Models},
  author = {Haozhe Huang and Yudong Xu and Abhijoy Mandal and Alán Aspuru-Guzik},
  year = {2026},
  abstract = {Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods are fundamentally limited by working on a per-sample basis, treating each generated state monolithical},
  url = {https://arxiv.org/abs/2607.11758},
  keywords = {cs.LG},
  eprint = {2607.11758},
  archiveprefix = {arXiv},
}

Metadata

{}