Paper Detail

MintAct: A Unified Visual Agent for Digital Environments

Mingfei Gao, Rui Tian, Haiming Gang, Bohan Zhai, Le Zhang, Yuanzheng Gong, Di Feng, Ege Özsoy, Kaixin Ma, Vishwesh Kirthivasan, Oğuzhan Fatih Kar, Roman Bachmann, Anders Boesen Lindbo Larsen, Afshin Dehghan

arxiv Score 15.8

Published 2026-09-18 · First seen 2026-09-21

Research Track B · General AI

Abstract

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{gao2026mintact,
  title = {MintAct: A Unified Visual Agent for Digital Environments},
  author = {Mingfei Gao and Rui Tian and Haiming Gang and Bohan Zhai and Le Zhang and Yuanzheng Gong and Di Feng and Ege Özsoy and Kaixin Ma and Vishwesh Kirthivasan and Oğuzhan Fatih Kar and Roman Bachmann and Anders Boesen Lindbo Larsen and Afshin Dehghan},
  year = {2026},
  abstract = {We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds},
  url = {https://arxiv.org/abs/2609.22083},
  keywords = {cs.CV},
  eprint = {2609.22083},
  archiveprefix = {arXiv},
}

Metadata

{}