Machine Learning-Based QSAR Classifications for PIM Kinases Inhibition Prediction: Towards the Neoplastic in Silico Drug Design
摘要
Promoting the use of strong AI tools in computational drug designing is a promising way to avoid early-stage failures of cancer drug discovery process. We build an inhibition targeted machine learning classifications, aiming to model the structure/activity relationships for PIM 1/2/3 protein kinases inhibitors, using different decision trees-based algorithms, starting from the data curation and analysis of previous experimental measurements. The therapeutic targets being studied are a family of serine/threonine protein kinases directly involved in various cellular processes, they have been implicated in cancer progression and identified as highly oncogenic. The constructed models showed Random Forest (RF) performances slightly better than XGBoost for the PIM 1 (+1% of difference in the accuracy scores), and XGBoost significant robustness for the PIM 2 and 3 datasets (+2% and + 4%, respectively), whereas the SVM algorithms were found to present a poor predictive ability from our datasets, either with a linear or a radial basis functional kernel. The benchmarking led to the selection of the strongest models: 85% of prediction accuracy for PIM 1 and PIM 2 datasets and 82% for the PIM 3 dataset. Data modeling along with technical methodology are discussed in details and the predictive strength of both RF and XGBoost algorithms on these data types is examined.