The accurate prediction of glass transition temperature ( \(T_g\) ) is crucial for materials design but often relies on melting point dependencies, limiting its applicability in inverse design problems. We present a data-driven approach leveraging machine learning (ML) and symbolic regression to predict \(T_g\) based solely on molecular structure. Using the BIMOG dataset, we extract key structural features—including molecular branching, computed from SMILES representations using RDKit and PySMILES, and atomic composition ratios (C, CH, O, etc.)—to enhance predictive accuracy. We apply multiple ML models, including Linear Regression, Random Forest, Gradient Boosting, XGBoost, and Extra Trees, achieving \(R^2\) scores comparable to traditional approaches that depend on melting point data. Finally, we employ genetic programming for symbolic regression to derive an interpretable equation for \(T_g\) . Our results demonstrate that incorporating structural descriptors allows for accurate and generalizable \(T_g\) prediction without requiring melting point information, making this method well-suited for inverse materials design. This work highlights how computational approaches can improve the tractability of complex materials problems, aligning with the broader goal of integrating physics-based and data-driven methods for materials discovery.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data-Driven Prediction of Glass Transition Temperature Using Molecular Structural Features

  • Sunny Kaushik,
  • Rohit Mogli,
  • Riddhika Mahalanabis,
  • Balakrishnan Ashok

摘要

The accurate prediction of glass transition temperature ( \(T_g\) ) is crucial for materials design but often relies on melting point dependencies, limiting its applicability in inverse design problems. We present a data-driven approach leveraging machine learning (ML) and symbolic regression to predict \(T_g\) based solely on molecular structure. Using the BIMOG dataset, we extract key structural features—including molecular branching, computed from SMILES representations using RDKit and PySMILES, and atomic composition ratios (C, CH, O, etc.)—to enhance predictive accuracy. We apply multiple ML models, including Linear Regression, Random Forest, Gradient Boosting, XGBoost, and Extra Trees, achieving \(R^2\) scores comparable to traditional approaches that depend on melting point data. Finally, we employ genetic programming for symbolic regression to derive an interpretable equation for \(T_g\) . Our results demonstrate that incorporating structural descriptors allows for accurate and generalizable \(T_g\) prediction without requiring melting point information, making this method well-suited for inverse materials design. This work highlights how computational approaches can improve the tractability of complex materials problems, aligning with the broader goal of integrating physics-based and data-driven methods for materials discovery.