<p>The identification of <b>protein target classes</b> is a key step in drug discovery, as it enables prioritization of screening campaigns and supports target-based drug repurpose. In this study, we developed a deep-learning pipeline based on a multilayer perceptron (MLP) trained on 15,804 curated compounds representing four major pharmacological target classes: G protein–coupled receptors (GPCRs), kinases, nuclear receptors, and transporters. Using extended connectivity fingerprints (ECFP4) as molecular descriptors, the model achieved 96% accuracy in internal cross-validation and 87% accuracy on an external test set, demonstrating performance comparable to ensemble classifiers such as Random Forest, XGBoost, and LightGBM. Class-specific F1 scores confirmed robust and balanced predictions across GPCR, kinase, nuclear receptor, and transporter categories. Model interpretability was addressed using SHAP values, which highlighted pharmacophore-like substructures consistent with known ligand–target interactions. Application to reference drugs further validated predictive utility, with correct assignment of most compounds to their canonical <b>protein target class</b>. The final MLP model was deployed as a user-friendly web application to facilitate accessible <b>protein class prediction</b> for novel compounds. Overall, this work presents a reliable and interpretable computational framework to support <b>target-class-based drug discovery and repositioning.</b></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DeepTargetClass: a web-based platform for predicting protein target classes of small molecules

  • Mebarka Ouassaf,
  • Bader Y. Alhatlani

摘要

The identification of protein target classes is a key step in drug discovery, as it enables prioritization of screening campaigns and supports target-based drug repurpose. In this study, we developed a deep-learning pipeline based on a multilayer perceptron (MLP) trained on 15,804 curated compounds representing four major pharmacological target classes: G protein–coupled receptors (GPCRs), kinases, nuclear receptors, and transporters. Using extended connectivity fingerprints (ECFP4) as molecular descriptors, the model achieved 96% accuracy in internal cross-validation and 87% accuracy on an external test set, demonstrating performance comparable to ensemble classifiers such as Random Forest, XGBoost, and LightGBM. Class-specific F1 scores confirmed robust and balanced predictions across GPCR, kinase, nuclear receptor, and transporter categories. Model interpretability was addressed using SHAP values, which highlighted pharmacophore-like substructures consistent with known ligand–target interactions. Application to reference drugs further validated predictive utility, with correct assignment of most compounds to their canonical protein target class. The final MLP model was deployed as a user-friendly web application to facilitate accessible protein class prediction for novel compounds. Overall, this work presents a reliable and interpretable computational framework to support target-class-based drug discovery and repositioning.