PaperScope
LIVE · 2026-09-29 05:40 UTC

One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion

Yulin Yuan, Ying Zhang, Xiangming Meng

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33698 v1
Category
Submitted
2026-09-27

Abstract

Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF's throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.

Comment: 25 pages, 7 figures, and 8 tables. Includes appendices

arXiv abs page · PDF · same-day batch