Paper Detail

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian Böther, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy

huggingface Score 10.5

Published 2026-06-26 · First seen 2026-07-06

General AI

Abstract

Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we collect 160 datasets spanning four data types -- image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data -- into a corpus of 6T multimodal tokens. DCVLM allows participants to test curation strategies (filtering, mixing, formatting, sampling) across 1B-8B models and 6.25B-200B token budgets. Models are then evaluated on a carefully selected suite of up to 52 downstream benchmarks across 9 domains. We conduct extensive experiments on DCVLM and find that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales. The resulting dataset, DCVLM-Baseline, enables training an 8B VLM to 63.6% accuracy on our 33-task core suite with 200B training tokens. Compared to FineVision, the state-of-the-art open VLM training dataset, this represents an improvement of +5.4pp. DCVLM and all accompanying artifacts will be made publicly available at https://www.datacomp.ai/dcvlm/.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{farina2026datacomp,
  title = {DataComp-VLM: Improved Open Datasets for Vision-Language Models},
  author = {Matteo Farina and Vishaal Udandarao and Thao Nguyen and Selim Kuzucu and Maximilian Böther and Andreas Hochlehnert and Adhiraj Ghosh and Marianna Nezhurina and Karsten Roth and Joschka Struber and Yuhui Zhang and Sebastian Dziadzio and Elaine Sui and Soumya Jahagirdar and Dhruba Ghosh and Hasan Hammoud and Thomas De Min and Simone Caldarella and Jehanzeb Mirza and Sedrick Keh and Mehdi Cherti and Hilde Kuehne and Bernt Schiele and Serena Yeung-Levy and Muhammad Ferjad Naeem and Federico Tombari and Ana Klimovic and Elisa Ricci and Matthias Bethge and Sewoong Oh and Ameya Prabhu and Alessio Tonioni and Jenia Jitsev and Massimiliano Mancini and Ludwig Schmidt and Nikhil Parthasarathy},
  year = {2026},
  abstract = {Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we collect 160 datasets spanning four data types -- image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data -- into a corpus of 6T },
  url = {https://huggingface.co/papers/2606.28551},
  keywords = {Vision-Language Models, data curation, multimodal tokens, data mixing, data filtering, downstream benchmarks, model scaling, code available, huggingface daily},
  eprint = {2606.28551},
  archiveprefix = {arXiv},
}

Metadata

{}