Protein-DNA Binding Sites Prediction via Integrating Pretrained Large Language Models and Contrastive Learning
摘要
Accurate identification of protein-DNA binding residues plays a vital role in elucidating molecular recognition and facilitating drug discovery. Nonetheless, conventional experimental techniques tend to be expensive, require substantial time, and demand considerable manual effort. The reliance of existing methods on handcrafted features often limits their ability to produce a high-quality initial representation. Although Protein Language Model (PLM) embeddings have proven effective, most methods depend on a single embedding source, limiting their generalizability and prediction accuracy. Here, we present IPDLPre, a novel model that combines an integrating pretrained PLM, a CNN-attention network, and contrastive learning for protein-DNA binding site prediction. Specifically, IPDLPre leverages three PLMs combined with a customized CNN-attention architecture to decode evolutionary features and generate residue-level binding confidence scores. To mitigate class imbalance, we employ a hybrid loss function combining triplet center loss and focal loss, enhancing feature representation and predictive performance. The experimental findings reveal that IPDLPre outperforms existing sequence-based techniques and delivers results comparable to those achieved by structure-based methods. Furthermore, IPDLPre demonstrates effective generalization on RNA-binding datasets, highlighting its broad applicability in predicting protein-nucleic acid interactions. In conclusion, IPDLPre provides a valuable strategy that holds great potential for advancing protein engineering and aiding drug development.
Graphical Abstract