Deep Learning of Protein–Ligand Interaction Prediction
摘要
Protein-small-molecule ligand interactions are central to structure-based drug design, where the core challenge lies in accurately characterizing binding strength and binding mechanisms. This work provides a systematic overview of key binding descriptors—binding constants (Kd/Ka, Ki, and IC50) and binding free energy (ΔG)—and explains how noncovalent interactions, including hydrogen bonding, salt bridges, hydrophobic effects, metal coordination, and cation–π interactions, collectively drive complex formation through structural and energetic complementarity. Methodologically, three mainstream routes for binding free energy/affinity estimation are summarized: physics-based molecular simulation and free-energy calculations (e.g., PMF, enhanced sampling, TI/FEP, MM/PBSA, and MM/GBSA), empirical scoring functions, and knowledge-based/statistical potential approaches. The review further highlights deep learning advances in structure-based affinity prediction and docking re-scoring, covering dataset construction (e.g., PDBbind and its benchmark subsets such as CASF), common molecular representations (molecular descriptors and fingerprints, 3D grid/voxel encodings), and the typical modeling workflow from training and hyperparameter tuning to evaluation. Representative models—including RF-Score, AtomNet, Pafnucy, Kdeep, OnionNet, and PLEC—are discussed for performance comparison. Finally, using feature extraction examples from DeepChem (grid features and fingerprints), this work illustrates random forest and multitask neural network modeling on PDBbind with correlation-based assessment, demonstrating the strong potential of deep learning to improve affinity prediction accuracy and accelerate virtual screening.