Background <p>The geo-traceability of cotton is crucial for ensuring the quality and integrity of cotton brands. However, effective methods for achieving this traceability are currently lacking. This study investigates the potential of explainable machine learning for the geo-traceability of raw cotton.</p> Results <p>The findings indicate that principal component analysis (PCA) exhibits limited effectiveness in tracing cotton origins. In contrast, partial least squares discriminant analysis (PLS-DA) demonstrates superior classification performance, identifying seven discriminating variables: Na, Mn, Ba, Rb, Al, As, and Pb. The use of decision tree (DT), support vector machine (SVM), and random forest (RF) models for origin discrimination yielded accuracies of 90%, 87%, and 97%, respectively. Notably, the light gradient boosting machine (LightGBM) model achieved&#xa0;perfect performance metrics,&#xa0;with accuracy, precision, and recall rate&#xa0;all reaching 100%&#xa0;on the test set. The output of the LightGBM model was further evaluated using the SHapley Additive exPlanation (SHAP) technique, which highlighted differences in the elemental composition of raw cotton from various countries. Specifically, the elements Pb, Ni, Na, Al, As, Ba, and Rb significantly influenced the model's predictions.</p> Conclusion <p>These findings suggest that explainable machine learning techniques can provide insights into the complex relationships between geographic information and raw cotton. Consequently, these methodologies enhances the precision and reliability of geographic traceability for raw cotton.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning improve the discrimination of raw cotton from different countries

  • Wang Tian,
  • Xu Shuangjiao,
  • Wei Jingyan,
  • Wang Ming,
  • Du Weidong,
  • Tian Xinquan,
  • Ma Lei

摘要

Background

The geo-traceability of cotton is crucial for ensuring the quality and integrity of cotton brands. However, effective methods for achieving this traceability are currently lacking. This study investigates the potential of explainable machine learning for the geo-traceability of raw cotton.

Results

The findings indicate that principal component analysis (PCA) exhibits limited effectiveness in tracing cotton origins. In contrast, partial least squares discriminant analysis (PLS-DA) demonstrates superior classification performance, identifying seven discriminating variables: Na, Mn, Ba, Rb, Al, As, and Pb. The use of decision tree (DT), support vector machine (SVM), and random forest (RF) models for origin discrimination yielded accuracies of 90%, 87%, and 97%, respectively. Notably, the light gradient boosting machine (LightGBM) model achieved perfect performance metrics, with accuracy, precision, and recall rate all reaching 100% on the test set. The output of the LightGBM model was further evaluated using the SHapley Additive exPlanation (SHAP) technique, which highlighted differences in the elemental composition of raw cotton from various countries. Specifically, the elements Pb, Ni, Na, Al, As, Ba, and Rb significantly influenced the model's predictions.

Conclusion

These findings suggest that explainable machine learning techniques can provide insights into the complex relationships between geographic information and raw cotton. Consequently, these methodologies enhances the precision and reliability of geographic traceability for raw cotton.