Paper Detail

A Distributional Optimisation Perspective on Combining Models in Deep Learning

Congye Wang, Yan Lin, Zheyang Shen, Matthew A. Fisher, Chris. J. Oates

arxiv Score 15.3

Published 2026-09-21 · First seen 2026-09-22

General AI

Abstract

Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen separately, and by ad hoc means. Recent advances in distributional optimisation (i.e. where the optimisation occurs over the set of probability distributions) offer an opportunity for principled joint training, viewing the collection of models as a discrete distribution whose support points are to be optimised, but the potential of these methods is not well-understood. In this paper we (1) cast two standard combination strategies - ensembles and low-rank adapter averaging - as entropy-regularised distributional optimisation, observing that the resulting objective is convex in the ensemble case but not in the adapter-averaging case, so that existing convergence guarantees for mean field Langevin dynamics transfer only to the former; (2) assess existing and novel algorithms for this task, including a functional variant of variational gradient descent; and (3) report an empirical study spanning synthetic classification tasks and fine-tuning of large language models on a commonsense reasoning benchmark.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{wang2026distributional,
  title = {A Distributional Optimisation Perspective on Combining Models in Deep Learning},
  author = {Congye Wang and Yan Lin and Zheyang Shen and Matthew A. Fisher and Chris. J. Oates},
  year = {2026},
  abstract = {Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen separately, and by ad hoc means. Recent advances in distributional optimisation (i.e. where the optimisation occurs over the set of probability distributions) offer an opportunity for principled joint training, viewing the collection of models as a discrete distribution whose support points are to be optimi},
  url = {https://arxiv.org/abs/2609.24328},
  keywords = {cs.LG, stat.ML},
  eprint = {2609.24328},
  archiveprefix = {arXiv},
}

Metadata

{}