<p>The development of novel alloys is limited by the need for resource-intensive experimentation. A large amount of context-dependent information is available in peer-reviewed research articles as unstructured text data, making it difficult for traditional machine learning algorithms to utilize it effectively. Existing transformer-based models such as AlloyBERT, MatBERT, polyBERT and MaterialBERT function on combined property targets which depends on DFT calculated descriptors. These models are restricted to single material classes and prioritize linguistic standards over mechanical property regression. To address these limitations, the current study proposes AliBERT (Alloy Lite BERT), a domain-specific and parameter-efficient transformer encoder. ALiBERT is utilized to predict yield strength (YS), ultimate tensile strength (UTS) and % elongation of MPEA directly form textual alloy composition descriptors without the need for pre-computed physical descriptors. ALiBERT uses cross-layer parameter sharing and factorized embedding to minimize the number of trainable parameters by 80% compared to full BERT. The models were pretrained and fine-tuned using the Citrine Informatics MPEA dataset, which contains 1545 experimentally described entries over 30 elements, domain-specific vocabulary of 1131 tokens and a masked language modeling objective. The fine-tuned ALiBERT embeddings were then used as input features for a hybrid ensemble RFXGB regressor. It combines random forest and XGBoost base learning with linear and ridge stacking. A fivefold cross-validation and Bayesian hyperparameter optimization were carried out. The ridge-stacked RFXGB model achieved RMSE value of 222.23&#xa0;MPa for UTS, 226.34&#xa0;MPa for YS and 228.78% elongation with R<sup>2</sup> values of 0.983, 0.970 and 0.968, respectively. The error reduction was almost 40% based on fine-tuning compared to baseline models. External validation confirmed composition predictive reliability by predicting ten out of distribution MPEA compositions from peer-reviewed literature. It achieved generalization R<sup>2</sup> values of 0.9704 for YS, 0.9228 for UTS and 0.9403 for % elongation. SHAP summary plots revealed that test temperature, niobium, molybednum and aluminum concentration are the factors influencing mechanical properties prediction of MPEA. These findings correlate with Peierls barrier theory, Fleischer–Labusch solid solution strengthening and B2 boundary hardening mechanism. In comparison with conventional ML models and domain-specific language models, the proposed ALiBERT—RFXGB framework introduces a new paradigm of alloy property prediction with a latency of less than 12&#xa0;ms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ALiBERT: Transformer-Based NLP and Hybrid Ensemble Learning Framework for Predicting Material Properties

  • M. Subramanian,
  • N. Arunkumar,
  • M. Diviya,
  • T. Balasubramanian

摘要

The development of novel alloys is limited by the need for resource-intensive experimentation. A large amount of context-dependent information is available in peer-reviewed research articles as unstructured text data, making it difficult for traditional machine learning algorithms to utilize it effectively. Existing transformer-based models such as AlloyBERT, MatBERT, polyBERT and MaterialBERT function on combined property targets which depends on DFT calculated descriptors. These models are restricted to single material classes and prioritize linguistic standards over mechanical property regression. To address these limitations, the current study proposes AliBERT (Alloy Lite BERT), a domain-specific and parameter-efficient transformer encoder. ALiBERT is utilized to predict yield strength (YS), ultimate tensile strength (UTS) and % elongation of MPEA directly form textual alloy composition descriptors without the need for pre-computed physical descriptors. ALiBERT uses cross-layer parameter sharing and factorized embedding to minimize the number of trainable parameters by 80% compared to full BERT. The models were pretrained and fine-tuned using the Citrine Informatics MPEA dataset, which contains 1545 experimentally described entries over 30 elements, domain-specific vocabulary of 1131 tokens and a masked language modeling objective. The fine-tuned ALiBERT embeddings were then used as input features for a hybrid ensemble RFXGB regressor. It combines random forest and XGBoost base learning with linear and ridge stacking. A fivefold cross-validation and Bayesian hyperparameter optimization were carried out. The ridge-stacked RFXGB model achieved RMSE value of 222.23 MPa for UTS, 226.34 MPa for YS and 228.78% elongation with R2 values of 0.983, 0.970 and 0.968, respectively. The error reduction was almost 40% based on fine-tuning compared to baseline models. External validation confirmed composition predictive reliability by predicting ten out of distribution MPEA compositions from peer-reviewed literature. It achieved generalization R2 values of 0.9704 for YS, 0.9228 for UTS and 0.9403 for % elongation. SHAP summary plots revealed that test temperature, niobium, molybednum and aluminum concentration are the factors influencing mechanical properties prediction of MPEA. These findings correlate with Peierls barrier theory, Fleischer–Labusch solid solution strengthening and B2 boundary hardening mechanism. In comparison with conventional ML models and domain-specific language models, the proposed ALiBERT—RFXGB framework introduces a new paradigm of alloy property prediction with a latency of less than 12 ms.