Discovery of novel High-Tc superconductors via machine learning-based random forest model
摘要
Superconductivity, a quantum mechanical marvel in condensed matter systems, continues to challenge fundamental understanding of composition-property relationships in superconducting materials. To address this knowledge gap, we present a machine learning framework analyzing 16,413 superconducting compounds from the SuperCon database. Using only chemical formulas as input, our random forest model achieves 93.5% accuracy and a relative root mean square error of 0.13 in predicting critical temperatures (Tc) through five-fold cross-validation. Additionally, the methodology demonstrates dual analytical capabilities through material-class-specific regression and Tc-range classification: Regression models tailored for iron-based (N = 1557), cuprate (N = 4403) and other (N = 6479) superconductors exhibit robust generalizability with R² >0.83 on test sets, while classification models categorizing materials into low- (< 10 K), medium- (10–77 K), and high-Tc (≥ 77 K) groups achieve 91% F1-score using random forest classifiers. When deployed for high-throughput screening of 4914 ABO3-type perovskites, the optimized model identifies three thermodynamically stable candidates (Ehull = 0 meV/atom) with predicted Tc >70 K. This data-driven framework not only deciphers composition-property relationships but also establishes an accelerated pathway for targeted discovery of high-temperature superconductors, bridging materials informatics with experimental synthesis priorities.