Machine Learning Approach Versus AutoML to Predict the Bioactivity of a Therapeutic Target Related to Cancer
摘要
This study presents a machine-learning approach to develop potent EGFR inhibitors for breast cancer treatment, leveraging data from the ChEMBL database and employing chemoinformatics and RDKit for model generation. Despite achieving high accuracy, the predictive model’s R2 score suggests room for improvement. The research compares traditional machine learning models (e.g., SVR, Random Forest, XGBoost, CatBoost) against automated machine learning (AutoML) tools like Pycaret and H2O AutoML, focusing on efficiency and performance metrics (RMSE, MAE, R2). Findings indicate that while the Random Forest model excels in traditional settings, AutoML offers significant time-saving advantages, albeit with slightly lower performance. This work underscores the potential of integrating AI into drug discovery, balancing between manual expertise in model building and the accelerated development process afforded by AutoML.