PaperScope
LIVE · 2026-09-29 05:40 UTC

Flow-Matching-Based Protein Structure Tokenizer Made Efficient and Easy

Zhe Zhang, Yikai Zhang, Jiangtao Feng, Ya-Qin Zhang, Wei-Ying Ma, Hao Zhou

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33129 v1
Category
Submitted
2026-09-27

Abstract

As the bridge between protein modality and discrete modeling, protein structure tokenization still largely relies on heavily engineered training objectives tailored to specific downstream tasks and large training datasets, which hinders its transfer to broader application scenarios. To address this issue, we propose ProFiT, a lightweight flow matching tokenizer. With simple training strategies that encourage healthy codebook utilization, ProFiT can be trained efficiently and naturally learns semantically meaningful representations without any manual semantic alignment, while achieving reconstruction quality and generalization that match or surpass those of substantially larger tokenizers. We conduct extensive evaluations across a wide range of settings and demonstrate that ProFiT is a plug-and-play tokenizer adaptable to diverse downstream tasks. This study further reveals the significant potential of the flow matching tokenizer paradigm. Our code is publicly available at https://github.com/QDKStorm/ProFiT.

arXiv abs page · PDF · same-day batch