<p>This study developed and validated Quantitative Structure-Activity Relationship models to predict the inhibitory activity (pIC<sub>50</sub>) of 225 EGFR inhibitors. A genetic algorithm selected eight molecular descriptors, which were used to construct two models: a multiple linear regression (MLR) and a stacked ensemble regression (SER). The SER model showed only marginally higher accuracy (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\Delta r^2 = +0.022\)</EquationSource> </InlineEquation>) but exhibited greater predictive instability (<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\Delta r^2_{m(test)} = 0.0802\)</EquationSource> </InlineEquation> vs. MLR’s 0.0184) and reduced interpretability. Thus, MLR was retained as the primary model due to its OECD-compliant mechanistic transparency and superior generalizability. Rigorous applicability domain analysis confirmed the MLR model’s reliability. Notably, molecular docking (PDB ID: 8A27) identified a top-ranked inhibitor (Compound 121) with high binding affinity (<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(-12.023\)</EquationSource> </InlineEquation> kcal/mol), forming critical hydrogen bonds and hydrophobic interactions with EGFR’s active site. Virtual screening of 32 structural analogs of Compound 121 revealed additional promising candidates. This work provides a robust framework for EGFR inhibitor discovery, combining computational modeling with structural insights.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Discovery of novel natural product-derived EGFR inhibitors using multiple linear regression, stacked ensemble regression, and fingerprinting approaches

  • Said Bitam,
  • Mabrouk Hamadache,
  • Salah Hanini

摘要

This study developed and validated Quantitative Structure-Activity Relationship models to predict the inhibitory activity (pIC50) of 225 EGFR inhibitors. A genetic algorithm selected eight molecular descriptors, which were used to construct two models: a multiple linear regression (MLR) and a stacked ensemble regression (SER). The SER model showed only marginally higher accuracy ( \(\Delta r^2 = +0.022\) ) but exhibited greater predictive instability ( \(\Delta r^2_{m(test)} = 0.0802\) vs. MLR’s 0.0184) and reduced interpretability. Thus, MLR was retained as the primary model due to its OECD-compliant mechanistic transparency and superior generalizability. Rigorous applicability domain analysis confirmed the MLR model’s reliability. Notably, molecular docking (PDB ID: 8A27) identified a top-ranked inhibitor (Compound 121) with high binding affinity ( \(-12.023\) kcal/mol), forming critical hydrogen bonds and hydrophobic interactions with EGFR’s active site. Virtual screening of 32 structural analogs of Compound 121 revealed additional promising candidates. This work provides a robust framework for EGFR inhibitor discovery, combining computational modeling with structural insights.