PaperScope
LIVE · 2026-10-05 05:40 UTC

Toward Omni Multimodal Graph Foundation Model: A Topology-Driven Binding Approach

Xunkai Li, Chenxi Wan, Yinlin Zhu, Wang Luo, Hongchao Qin, Rong-Hua Li, Guoren Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.02881 v1
Category
Submitted
2026-10-02

Abstract

Multimodal graph foundation models (MGFMs) seek to learn generalizable representations from large-scale graphs with heterogeneous node modalities. However, real-world Multimodal-Attributed Graphs (MAGs) often contain incomplete node attributes, limiting the scale and diversity of available pretraining corpora. Besides, existing MGFMs primarily incorporate graph topology as structural context, overlooking its role in guiding multimodal binding and shaping a unified representation space. To address these challenges, we propose GraphBind, a topology-driven approach that uses graph topology to bind rich modality information into a unified shared space. GraphBind is motivated by the stability of graph topology, which provides structural references and complementary semantic information for multimodal binding. Concretely, GraphBind uses topology to organize self semantics and reliable neighborhood semantics into a global shared space that integrates structure and semantics, and adapts this space to discriminative and generative tasks through lightweight interfaces. Extensive experiments against 11 representative baselines demonstrate that GraphBind achieves leading performance on both discriminative and generative tasks, with relative improvements of up to 28.1% over the strongest baseline.

arXiv abs page · PDF · same-day batch