<p><i>Pericarpium Citri Reticulatae</i> (Chenpi) is valued for its age-dependent quality, yet market fraud necessitates objective identification methods. This study established a multi-source data fusion approach combining Fourier transform infrared (FTIR) spectroscopy and gas chromatography-mass spectrometry (GC-MS) metabolomics with machine learning for Chenpi age classification. Sixty samples spanning three age groups (5, 10, and 15 years) were analyzed. For FTIR data, Savitzky-Golay smoothing combined with multiplicative scatter correction (SG + MSC) was identified as the optimal preprocessing method (silhouette score: 0.7417; CH index: 824.89). For GC-MS data, Log2 transformation yielded the best group separation (score: 2.62) and identified 597 differentially abundant compounds via F-test. Three classification models were evaluated using 5-fold stratified cross-validation. The mid-level data fusion combined with support vector machine (SVM) achieved the highest accuracy of 96.7%, outperforming single-source models. These results have demonstrated that multi-source data fusion significantly enhanced Chenpi age classification, providing a robust approach for Chenpi authentication.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-source data feature fusion and machine learning algorithms for storage age identification of Pericarpium Citri Reticulatae (Chenpi)

  • Songmei Wang,
  • Yifan Tu,
  • Wenxuan Deng,
  • Zhanming Li,
  • Ming Li

摘要

Pericarpium Citri Reticulatae (Chenpi) is valued for its age-dependent quality, yet market fraud necessitates objective identification methods. This study established a multi-source data fusion approach combining Fourier transform infrared (FTIR) spectroscopy and gas chromatography-mass spectrometry (GC-MS) metabolomics with machine learning for Chenpi age classification. Sixty samples spanning three age groups (5, 10, and 15 years) were analyzed. For FTIR data, Savitzky-Golay smoothing combined with multiplicative scatter correction (SG + MSC) was identified as the optimal preprocessing method (silhouette score: 0.7417; CH index: 824.89). For GC-MS data, Log2 transformation yielded the best group separation (score: 2.62) and identified 597 differentially abundant compounds via F-test. Three classification models were evaluated using 5-fold stratified cross-validation. The mid-level data fusion combined with support vector machine (SVM) achieved the highest accuracy of 96.7%, outperforming single-source models. These results have demonstrated that multi-source data fusion significantly enhanced Chenpi age classification, providing a robust approach for Chenpi authentication.