Paper Detail

Form and Void: Entangled Composition through an Autonomous AI Agent

Shiwen Wang, Jian Yang, Xu Wang, Xincan Wang, Weiming Dong

arxiv Score 15.6

Published 2026-10-01 · First seen 2026-10-02

General AI

Abstract

Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbf{F}orm \textbf{a}nd \textbf{V}oid \textbf{A}gent (\textbf{FaV-A}), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{wang2026form,
  title = {Form and Void: Entangled Composition through an Autonomous AI Agent},
  author = {Shiwen Wang and Jian Yang and Xu Wang and Xincan Wang and Weiming Dong},
  year = {2026},
  abstract = {Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains diffic},
  url = {https://arxiv.org/abs/2610.02045},
  keywords = {cs.CV, cs.MA},
  eprint = {2610.02045},
  archiveprefix = {arXiv},
}

Metadata

{}