PaperScope
LIVE · 2026-10-06 05:40 UTC

On the Geometry of Multimodal Saturation: Riemannian VICReg

Nessim Ben Abbes, Duc Han Le, Sabri Mtibaa, Van-Tam Nguyen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06096 v1
Category
Submitted
2026-10-05

Abstract

In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired runs under VICReg. The same failure occurs in 51.1% of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Riemannian VICReg (R-VICReg) generalizes classical VICReg: it aligns views by squared geodesic distance on learnable negative-curvature product factors and recovers VICReg exactly as curvature vanishes. Over the same 45 paired runs, R-VICReg raises the probability that the third modality helps from 44.4% to 64.4%, with gains concentrated where VICReg saturates.

Comment: 17 pages, 6 figures. Accepted to the Proceedings Track of the NeurIPS 2026 Workshop on Symmetry and Geometry in Neural Representations (NeurReps)

arXiv abs page · PDF · same-day batch