CoBdock-2: enhancing blind docking performance through hybrid feature selection combining ensemble and multimodel feature selection approaches
摘要
Identifying orthosteric binding sites and predicting small molecule affinities remains a key challenge in virtual screening. While blind docking explores the entire protein surface, its precision is hindered by the vast search space. Cavity detection-guided docking improves accuracy by narrowing focus to predicted pockets, but its effectiveness depends heavily on the quality of cavity detection tools. To overcome these limitations, we developed Consensus Blind Dock (CoBDock), a machine learning-based blind docking method that integrates molecular docking and cavity detection results to enhance binding site and pose prediction. Building on this, CoBDock-2 replaces traditional docking tools by extracting 1D numerical representations from protein, ligand, and interaction structural features, and applying advanced ensemble feature selection techniques. By evaluating 21 feature selection methods across 9,598 features, CoBDock-2 identifies key molecular characteristics of orthosteric binding sites. CoBDock-2 demonstrates consistent improvements over the original CoBDock across benchmark datasets (PDBBind v2020-general, MTi, ADS, DUD-E, CASF-2016), achieving 77% binding site identification accuracy (within 8 Å), 55% ligand pose prediction accuracy (RMSD