PaperScope
LIVE · 2026-09-09 05:40 UTC

Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist

Ming Cheng, Jiaying Gong, Hoda Eldardiry

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06406 v1
Category
Submitted
2026-09-06

Abstract

Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate on model parameters while overlooking the geometric structure of latent representation distributions across domains and tasks. To address these limitations, we propose Hierarchical Wasserstein Merging (HWM), a representation-level framework that models each domain-task specialist as a distribution of hidden representations on a shared support. HWM constructs task-level and global Wasserstein barycenters to capture within-task domain variation and cross-task structure, enabling either training-free specialist aggregation by Wasserstein-derived weights or training-based generalist learning through a hybrid Wasserstein alignment loss. Experiments on four NLP tasks across four domains per task show that HWM achieves superior effectiveness and generalization capability in MD-MTL settings.

Comment: 20 pages, 2 figures, accepted for publication in EMNLP 2026 Findings

arXiv abs page · PDF · same-day batch