Paper Detail

Embedding Prediction Helps Image Generation

Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu

arxiv Score 7.6

Published 2026-10-01 · First seen 2026-10-02

General AI

Abstract

In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{xu2026embedding,
  title = {Embedding Prediction Helps Image Generation},
  author = {Sihan Xu and Ji Xie and Zilin Wang and Hui Shen and Stella X. Yu},
  year = {2026},
  abstract = {In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to pred},
  url = {https://arxiv.org/abs/2610.02203},
  keywords = {cs.CV, cs.LG, Embedding, Computer science, Autoregressive model, Artificial intelligence, Transformer},
  eprint = {2610.02203},
  archiveprefix = {arXiv},
}

Metadata

{}