Paper Detail
Dang Quang Thien Tran, Quang V. Dang, Vinamra Tyagi, Sai Soorya Rao Veeravalli, Trang Nguyen, Ryan A. Rossi, Franck Dernoncourt, Nedim Lipka, Koustava Goswami, Samyadeep Basu
As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and model safety. While unimodal attributions have been explored in depth, the multimodal setting remains relatively under-researched. As a result, we introduce MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a document. To establish baseline results for the method, we introduce MultAttrEval, a complementary benchmark dataset annotated with fine-grained, ground-truth attributions for answer components grounded in multimodal source documents. To our knowledge, this is the first evaluation dataset designed specifically for multimodal attribution in long-form documents. Experimental results show that MultAttnAttrib consistently outperforms a variety of attribution-generation methods, including several strong prompting-based approaches and matches the latest frontier models such as GPT 5.4. Our method not only substantially improves attribution accuracy for both unimodal and multimodal attribution types, but also produces attributions at up to one-seventh of the direct inference latency compared to prompting on the same base model.
No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.
No ranking explanation is available yet.
No tags.
@misc{tran2026multattnattrib,
title = {MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering},
author = {Dang Quang Thien Tran and Quang V. Dang and Vinamra Tyagi and Sai Soorya Rao Veeravalli and Trang Nguyen and Ryan A. Rossi and Franck Dernoncourt and Nedim Lipka and Koustava Goswami and Samyadeep Basu},
year = {2026},
abstract = {As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and model safety. While unimodal attributions have been explored in depth, the multimodal setting remains relatively under-researched. As a result, we introduce MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a docu},
url = {https://huggingface.co/papers/2607.01420},
keywords = {multimodal attribution, attention heads, calibrated thresholds, grounding QA systems, attribution-generation method, prefill pass, multimodal source documents, long-form documents, attribution accuracy, inference latency, huggingface daily},
eprint = {2607.01420},
archiveprefix = {arXiv},
}
{}