PaperScope
LIVE · 2026-09-29 05:40 UTC

BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions

Thanina Hamitouch, Khadidja Henni, Abdelkrim Arie, Amina Selma Haichour, Neila Mezghani, Lina Abou-Abbas

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33254 v1
Category
Submitted
2026-09-27

Abstract

Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce and unevenly distributed, large pretrained chemical and protein encoders are expensive to fine-tune end-to-end, and independently encoded sequences do not capture pair-specific dependencies. We present BERT4DTI, which encodes SMILES strings with ChemBERTa and amino-acid sequences with ProtBERT, applies bidirectional mutual attention between token-level representations, and classifies the resulting interaction features using convolutional layers and a multilayer perceptron. To reduce trainable size, ProtBERT is truncated to 18 retained layers and only the last two layers of each encoder are fine-tuned. On BIOSNAP, DAVIS and BindingDB, BERT4DTI is competitive, achieving the best ROC-AUC and PR-AUC on BIOSNAP and the highest sensitivity on all three benchmarks. An ablation on DAVIS shows that mutual attention improves PR-AUC and specificity. With 125M trainable parameters compared with 353M for full BERT fine-tuning, BERT4DTI provides a favourable performance-parameter trade-off for sequence-based DTI screening, while leaving runtime profiling, calibration and leakage-audited validation for future work.

Comment: Accepted at CIKM 2026

arXiv abs page · PDF · same-day batch