Paper Detail

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

Enjun Du, Siyi Liu, Zirong Chen, Xinyu Zuo, Jinwen Luo, Ruiwen Tao, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang

huggingface Score 13.5

Published 2026-08-21 · First seen 2026-08-24

General AI

Abstract

Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{du2026evirank,
  title = {EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking},
  author = {Enjun Du and Siyi Liu and Zirong Chen and Xinyu Zuo and Jinwen Luo and Ruiwen Tao and Lisheng Duan and Haijin Liang and Jin Ma and Junfu Pu and Yongqi Zhang},
  year = {2026},
  abstract = {Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction p},
  url = {https://huggingface.co/papers/2608.20886},
  keywords = {multimodal image re-ranking, compositional queries, semantic constraint satisfaction, evidence package, rubric scoring, listwise comparison, training-free, distilled student, huggingface daily},
  eprint = {2608.20886},
  archiveprefix = {arXiv},
}

Metadata

{}