<p>There is a growing need for industry and global regulatory agencies to develop rapid chemical safety assessment through more reliable theoretical models. Thus<i>,</i> quantitative structure–toxicity relationship (QSTR) models are preferred by regulators to bring chemicals to market rather than long and expensive animal testing. In this study, we evaluated four binary classification machine learning (ML) models (support vector machine, <i>k</i>-nearest neighbor, CART decision tree and random forest) for their ability to predict toxicity towards <i>Tetrahymena pyriformis</i> using 1416 benzene-derived compounds (749 chemicals evaluated and 697 synthetic toxicants) classified into two groups: non-toxic molecules (NTox) with 708 observations and toxic molecules (Tox) with 708 observations. Here, ML models have been developed on the basis of data mining methods using the <i>ClustOfvar</i> algorithm for optimal feature selection and SMOTE methods for data balancing, forgoing the hyperparameter tuning techniques of the statistical learning models used. Of the four ML models based on the results of the external validation set centered on fivefold cross-validation, the robust and explanatory CART-decision tree (DT) model achieved the best results (<i>Q</i> = 95.42%, <i>Pr</i> = 96.60%, <i>Re</i> = 94.67%, F_score = 95.62%, <i>Sp</i> = 96.27%, <i>MCC</i> = 0.91, and <i>AUC</i> = 1.0). Thus, a set of 10 decision rules for predicting BZC (benzene-derived compounds) toxicity, easy to understand by humans, was also identified. The methodologies proposed in this paper would be useful for QSTR modeling by filling data gaps, prioritizing, and focusing experiments on the most hazardous organic chemicals.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pollution risk assessment by designing predictive binary classification models of substituted benzenes centered on data mining and machine learning techniques

  • Aubin N’guessan,
  • Brice Dali,
  • Elvice Akori Esmel,
  • Logbo Mathias Moussé,
  • Nahossé Ziao,
  • Raymond Kré N’guessan,
  • Eugene Megnassan

摘要

There is a growing need for industry and global regulatory agencies to develop rapid chemical safety assessment through more reliable theoretical models. Thus, quantitative structure–toxicity relationship (QSTR) models are preferred by regulators to bring chemicals to market rather than long and expensive animal testing. In this study, we evaluated four binary classification machine learning (ML) models (support vector machine, k-nearest neighbor, CART decision tree and random forest) for their ability to predict toxicity towards Tetrahymena pyriformis using 1416 benzene-derived compounds (749 chemicals evaluated and 697 synthetic toxicants) classified into two groups: non-toxic molecules (NTox) with 708 observations and toxic molecules (Tox) with 708 observations. Here, ML models have been developed on the basis of data mining methods using the ClustOfvar algorithm for optimal feature selection and SMOTE methods for data balancing, forgoing the hyperparameter tuning techniques of the statistical learning models used. Of the four ML models based on the results of the external validation set centered on fivefold cross-validation, the robust and explanatory CART-decision tree (DT) model achieved the best results (Q = 95.42%, Pr = 96.60%, Re = 94.67%, F_score = 95.62%, Sp = 96.27%, MCC = 0.91, and AUC = 1.0). Thus, a set of 10 decision rules for predicting BZC (benzene-derived compounds) toxicity, easy to understand by humans, was also identified. The methodologies proposed in this paper would be useful for QSTR modeling by filling data gaps, prioritizing, and focusing experiments on the most hazardous organic chemicals.