<p>This study presents a rigorous machine-learning framework for modeling the cross-sections (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\sigma \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>σ</mi> </math></EquationSource> </InlineEquation>) of (<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(n, d\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>n</mi> <mo>,</mo> <mi>d</mi> </mrow> </math></EquationSource> </InlineEquation>) nuclear reactions at a standard fast-neutron energy of approximately 14.5&#xa0;MeV derived from a Deuterium–Tritium (D–T) fusion source. Given that the incident energy remains constant across the curated experimental dataset of 28 nuclides (<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(Z = 2-73\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>Z</mi> <mo>=</mo> <mn>2</mn> <mo>-</mo> <mn>73</mn> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(A = 3-181\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>A</mi> <mo>=</mo> <mn>3</mn> <mo>-</mo> <mn>181</mn> </mrow> </math></EquationSource> </InlineEquation>), the modeling is executed on a base-10 logarithmic scale to effectively accommodate cross-sections spanning nearly three orders of magnitude (<InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(0.15-66\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>0.15</mn> <mo>-</mo> <mn>66</mn> </mrow> </math></EquationSource> </InlineEquation>&#xa0;mb). To overcome the inherent constraints of the small-data regime, the regression-adapted Synthetic Minority Oversampling Technique (SMOTER) was implemented. Methodological integrity was preserved via a robust, leakage-free Leave-One-Out cross-validation (LOOCV) protocol, wherein the held-out validation target was restricted to real experimental nuclides, and SMOTER was applied exclusively within each training fold. Furthermore, the synthetic instances are physically constrained to integer-valued atomic numbers (<InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(Z\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>Z</mi> </math></EquationSource> </InlineEquation>) and mass numbers (<InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(A\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>A</mi> </math></EquationSource> </InlineEquation>) localized within the valley of <InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation>-stability. Employing a comprehensive suite of physically motivated nuclear structure descriptors, we systematically evaluated seven regressors (Lasso, Ridge, Elastic Net, Random Forest, Gradient Boosting, XGBoost, and Support Vector Regression) along with two meta-ensembles (Voting and Stacking). Under strict leakage-free evaluation, the Elastic Net model exhibited superior generalization capacity (R<sup>2</sup> = 0.816, 95% CI [0.537, 0.894], RMSE = 0.259, and a within-factor-of-two agreement of 78.6%), closely followed by Lasso (R<sup>2</sup> = 0.791), whereas tree- and kernel-based architectures performed poorly (e.g., XGBoost R<sup>2</sup> = 0.608). Notably, our optimized Elastic Net model outperformed the TALYS-based TENDL-2019 evaluation (R<sup>2</sup> = 0.686) for the same nuclide set. Finally, split conformal prediction intervals provide rigorously calibrated uncertainty quantification, yielding 89.3% empirical coverage against a 90% nominal confidence level. These findings underscore that in data-scarce nuclear modeling regimes, explicit regularization and physically informed feature engineering are far more decisive for model generalization than the nominal algorithmic complexity.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting (n, d) reaction cross-sections using regularized linear models and ensemble methods: a comparative machine learning study

  • Cafer Mert Yeşilkanat,
  • Fatiha Kadem,
  • Serkan Akkoyun,
  • Mohamed Belgaid

摘要

This study presents a rigorous machine-learning framework for modeling the cross-sections ( \(\sigma \) σ ) of ( \(n, d\) n , d ) nuclear reactions at a standard fast-neutron energy of approximately 14.5 MeV derived from a Deuterium–Tritium (D–T) fusion source. Given that the incident energy remains constant across the curated experimental dataset of 28 nuclides ( \(Z = 2-73\) Z = 2 - 73 , \(A = 3-181\) A = 3 - 181 ), the modeling is executed on a base-10 logarithmic scale to effectively accommodate cross-sections spanning nearly three orders of magnitude ( \(0.15-66\) 0.15 - 66  mb). To overcome the inherent constraints of the small-data regime, the regression-adapted Synthetic Minority Oversampling Technique (SMOTER) was implemented. Methodological integrity was preserved via a robust, leakage-free Leave-One-Out cross-validation (LOOCV) protocol, wherein the held-out validation target was restricted to real experimental nuclides, and SMOTER was applied exclusively within each training fold. Furthermore, the synthetic instances are physically constrained to integer-valued atomic numbers ( \(Z\) Z ) and mass numbers ( \(A\) A ) localized within the valley of \(\beta \) β -stability. Employing a comprehensive suite of physically motivated nuclear structure descriptors, we systematically evaluated seven regressors (Lasso, Ridge, Elastic Net, Random Forest, Gradient Boosting, XGBoost, and Support Vector Regression) along with two meta-ensembles (Voting and Stacking). Under strict leakage-free evaluation, the Elastic Net model exhibited superior generalization capacity (R2 = 0.816, 95% CI [0.537, 0.894], RMSE = 0.259, and a within-factor-of-two agreement of 78.6%), closely followed by Lasso (R2 = 0.791), whereas tree- and kernel-based architectures performed poorly (e.g., XGBoost R2 = 0.608). Notably, our optimized Elastic Net model outperformed the TALYS-based TENDL-2019 evaluation (R2 = 0.686) for the same nuclide set. Finally, split conformal prediction intervals provide rigorously calibrated uncertainty quantification, yielding 89.3% empirical coverage against a 90% nominal confidence level. These findings underscore that in data-scarce nuclear modeling regimes, explicit regularization and physically informed feature engineering are far more decisive for model generalization than the nominal algorithmic complexity.