<p>This study constructed multiple regression models to predict the flash points of structurally diverse organic hydrocarbons, including alkanes, alkenes, alkynes, and their mixtures. It systematically compared the performance of traditional machine learning methods (e.g., KNN, LASSO, SVR, DT, and AdaBoost), ensemble models (e.g., RF, XGBoost, and CatBoost), and deep neural networks (e.g., ANN, DNNA, and DNNB) under varying structural complexities. The selected molecular descriptors encompassed multiple dimensions such as topology, autocorrelation, charge, and electronegativity. In the modeling of the entire dataset, the Bayesian-optimized deep neural network (DNNB) performed best, achieving an <i>R</i><sup>2</sup> of 0.861 on the prediction set, an overall <i>R</i><sup>2</sup> of 0.918, and an MAPE of only 3.544%, demonstrating excellent stability and generalization ability. In contrast, traditional models such as KNN, RF, and XGBoost performed well on certain subsets (e.g., XGBoost reached an <i>R</i><sup>2</sup> of 0.972 on alkanes) but exhibited significantly decreased accuracy on more structurally complex datasets (e.g., SVR achieved only 0.649 <i>R</i><sup>2</sup> on the mixed set). ANN showed limited fitting ability on simple data (lowest <i>R</i><sup>2</sup> of 0.28) but its prediction accuracy improved markedly to 0.765 after multi-structure fusion, indicating strong learning capacity for complex structural data. Feature importance analysis further indicated structure-specific descriptor dependencies among different hydrocarbon categories: Alkanes were sensitive to global topological features such as AATS8m and ATSC0s; alkenes relied more on charge and geometric distribution features (e.g., MATS4c and GATS5c); while alkynes exhibited very low correlation, resulting in regression difficulties that necessitated hybrid modeling to enhance learning efficiency. Within the fused models, features such as SM1_DzZ (atomic block distance index), BIC1 (basic information content), and MATS2e (electronegativity-weighted Moran autocorrelation) emerged as core variables influencing flash point prediction. This study demonstrated that deep learning-based modeling strategies, especially the Bayesian-optimized DNNB model, could fully exploit nonlinear relationships between complex molecular structures and thermophysical properties, providing a more accurate and generalizable technical approach for flash point prediction in high-throughput screening and inherent safety design.</p> Graphical abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AI-based accurate and efficient flash point prediction for structurally diverse hydrocarbons via Bayesian-optimized deep neural networks

  • Fanzhi Meng,
  • Wei Xu,
  • Yanan Qian,
  • Feng Sun,
  • Bing Sun,
  • Zhe Yang

摘要

This study constructed multiple regression models to predict the flash points of structurally diverse organic hydrocarbons, including alkanes, alkenes, alkynes, and their mixtures. It systematically compared the performance of traditional machine learning methods (e.g., KNN, LASSO, SVR, DT, and AdaBoost), ensemble models (e.g., RF, XGBoost, and CatBoost), and deep neural networks (e.g., ANN, DNNA, and DNNB) under varying structural complexities. The selected molecular descriptors encompassed multiple dimensions such as topology, autocorrelation, charge, and electronegativity. In the modeling of the entire dataset, the Bayesian-optimized deep neural network (DNNB) performed best, achieving an R2 of 0.861 on the prediction set, an overall R2 of 0.918, and an MAPE of only 3.544%, demonstrating excellent stability and generalization ability. In contrast, traditional models such as KNN, RF, and XGBoost performed well on certain subsets (e.g., XGBoost reached an R2 of 0.972 on alkanes) but exhibited significantly decreased accuracy on more structurally complex datasets (e.g., SVR achieved only 0.649 R2 on the mixed set). ANN showed limited fitting ability on simple data (lowest R2 of 0.28) but its prediction accuracy improved markedly to 0.765 after multi-structure fusion, indicating strong learning capacity for complex structural data. Feature importance analysis further indicated structure-specific descriptor dependencies among different hydrocarbon categories: Alkanes were sensitive to global topological features such as AATS8m and ATSC0s; alkenes relied more on charge and geometric distribution features (e.g., MATS4c and GATS5c); while alkynes exhibited very low correlation, resulting in regression difficulties that necessitated hybrid modeling to enhance learning efficiency. Within the fused models, features such as SM1_DzZ (atomic block distance index), BIC1 (basic information content), and MATS2e (electronegativity-weighted Moran autocorrelation) emerged as core variables influencing flash point prediction. This study demonstrated that deep learning-based modeling strategies, especially the Bayesian-optimized DNNB model, could fully exploit nonlinear relationships between complex molecular structures and thermophysical properties, providing a more accurate and generalizable technical approach for flash point prediction in high-throughput screening and inherent safety design.

Graphical abstract