PaperScope
LIVE · 2026-09-09 05:40 UTC

Foundation Models for Generalizable Semantic and Goal-Oriented Communication

Boliang Liu, Wint Yi Poe, Riccardo Trivisonno, Giuseppe Caire

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07853 v1
Submitted
2026-09-07

Abstract

Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data remains a key weakness under tight rate budgets. Many existing systems overfit their training data and degrade sharply at very low bit rates because they attempt to compress the entire signal. We introduce Foundation Model-Guided Semantic and Goal-Oriented Communication (FMSGOC), a framework that uses broad visual-linguistic Foundation Model priors to mitigate overfitting. It further improves rate efficiency by concentrating bits on sparse, goal-aligned anchors and relying on generative foundation-model priors to reconstruct the masked regions. By decoupling what to send from how to reconstruct, a vision-language foundation model selects and transmits a sparse set of semantic anchors, while a pretrained diffusion model, fine-tuned for masked completion, reconstructs the image at the receiver. In our experiments, FMSGOC reaches 0.039 bits per pixel (BPP), maintains high semantic fidelity (cosine similarity 0.87-0.90 on CIFAR-10), remains robust on previously unseen inputs (0.83-0.86 on ImageNet), and shows good perceptual similarity (0.1278/0.1558, CIFAR-10/ImageNet), outperforming strong end-to-end baselines at lower bit rates.

Comment: 6 pages, IEEE ICC 2026

Journal: 2026 IEEE International Conference on Communications (ICC), pp. 1-6, 2026

arXiv abs page · PDF · same-day batch