CORTEXA
← Browse
arxivstat.MLcs.LG2026-07-05

Optimal Mixture-of-Experts Model Averaging for Conditional Generative Models

Shijin Gong, Baihua He, Xinyu Zhang

Conditional generative models have emerged as powerful tools for sampling from target conditional distributions, driving substantial advances across a wide range of scientific and applied domains. As these models proliferate, practitioners often face multiple plausible generators whose performance can vary with the task, data, or input condition. We propose an optimal model averaging framework for conditional generative models, allowing candidate generators to be combined even when they are accessible only through conditional samples without tractable densities. Specifically, we use a sample-based maximum mean discrepancy between conditional distributions, which first leads to a static model averaging method, StaticMA, assigning fixed weights to different candidates. In addition, we develop MoEMA (mixture-of-experts model averaging), an input-adaptive method that parameterizes covariate-dependent weights through a softmax neural-network gate. We establish in-sample and out-of-sample asymptotic optimality for the proposed methods, together with consistency of the estimated adaptive weight function under regularity conditions. The framework applies directly to Euclidean responses and extends to unstructured data by combining our formulation with fixed representation maps. Across a broad set of simulations and real-data studies spanning tabular, image, and text modalities, MoEMA generally improves over competing baselines, demonstrating the effectiveness of our proposed methods.

View free PDFSource page

Related papers

arxivcs.LGcs.AIstat.ML2026-07-13

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

Nikita Kozodoi, Zainab Afolabi, Jack Butler

Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically studying how training duration of d…

View free PDFSource page
arxivstat.MLcs.LG2026-07-18

Semi-Supervised Conditional Generative Learning through Stochastic Interpolation and Sufficient Representations

Changyu Liu, Yuling Jiao, Jian Huang

Conditional generative modeling remains a challenging problem in semi-supervised settings where labeled data is scarce but unlabeled samples are abundant. To effectively leverage structural information embedded within the unlabeled dataset and compensate for sparse conditioning s…

View free PDFSource page
arxivstat.MLcs.LG2026-07-06

Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models

Stanislas Strasman, Sobihan Surendran, Sylvain Le Corff

Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training re…

View free PDFSource page
arxivcs.LGmath.PRstat.ML2026-07-16

Diffusion models recover accurate mixture weights despite score function insensitivity

Andrew Dennehy, Ramchandran Muthukumar, Rebecca Willett, Nisha Chandramoorthy

Score-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture weights. We resolve this apparent paradox by rel…

View free PDFSource page