Paper Detail

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu

arxiv Score 12.2

Published 2026-08-06 · First seen 2026-08-07

General AI

Abstract

The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{wang2026illusion,
  title = {The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images},
  author = {Zhiheng Wang and Bo Peng and Lai Wei and Chaochao Lu},
  year = {2026},
  abstract = {The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a c},
  url = {https://arxiv.org/abs/2608.06270},
  keywords = {cs.AI, multimodal LLMs, visual tool-use, causal graph, observation-mediated paths, action-induced shortcuts, Visual Evidence Gain, policy miscalibration, Calling Without Looking, Looking Without Planning, illusion of visual tool-use, code available, huggingface daily},
  eprint = {2608.06270},
  archiveprefix = {arXiv},
}

Metadata

{}