PaperScope
LIVE · 2026-09-22 05:40 UTC

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23073 v1
Category
Submitted
2026-09-19

Abstract

Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.

Comment: Accepted to EMNLP 2026 Main Conference

arXiv abs page · PDF · same-day batch