Paper Detail

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Zirong Chen, Fuda Ye, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang

huggingface Score 13.5

Published 2026-08-30 · First seen 2026-09-03

General AI

Abstract

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{chen2026snapbench,
  title = {SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions},
  author = {Zirong Chen and Fuda Ye and Kuan Zhang and Enjun Du and Junfu Pu and Xinlei Wang and Xinyu Zuo and Lisheng Duan and Jin Ma and Yongqi Zhang},
  year = {2026},
  abstract = {Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 que},
  url = {https://huggingface.co/papers/2608.29607},
  keywords = {snap-and-ask retrieval, multimodal retrieval, dual-tower encoders, embedding-based VLMs, cross-modal fallback, MOOR, Modality-anchored Outlier-aware Optimal Reweighting, reliability-aware modality calibration, huggingface daily},
  eprint = {2608.29607},
  archiveprefix = {arXiv},
}

Metadata

{}