Paper Detail

Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari

arxiv Score 7.6

Published 2026-10-01 · First seen 2026-10-02

General AI

Abstract

Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote sensing tasks, ExpertLens matches or surpasses full fine-tuning while updating only 21.7 - 47.0% of model parameters and achieving a 4.0x average training speedup, and outperforms LoRA in both adaptation performance and training efficiency. These results show that sparsity introduced for efficiency can give rise to semantic modularity that is directly useful for efficient adaptation.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{marsili2026harnessing,
  title = {Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation},
  author = {Damiano Marsili and Raphi Kang and Aditya Mehta and Pietro Perona and Georgia Gkioxari},
  year = {2026},
  abstract = {Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-sp},
  url = {https://arxiv.org/abs/2610.02123},
  keywords = {cs.CV},
  eprint = {2610.02123},
  archiveprefix = {arXiv},
}

Metadata

{}