A method based on local linear reconstruction combined with spectral information entropy for constructing and updating apple origin discrimination models
摘要
When using near-infrared (NIR) techniques combined with machine learning to identify the origin of apples, the construction and updating of discrimination models is essential. However, when constructing and updating the models, due to the high dimensionality of NIR data and the redundancy of the updated sample set can lead to irrational sample partitioning and the selected updated samples are not effective enough, which in turn affects the prediction performance of the NIR models. Therefore, in this study, local linear reconstruction combined with spectral information entropy (LLR–SIE) as the sample selection method is proposed. The two traditional methods, the Kennard-Stone (KS) and the sample set portioning based on joint x–y distance (SPXY) are used to compare and verify the superiority of the proposed method. Discrimination models was constructed after dividing the sample set using LLR–SIE method, KS method and SPXY method, and the results of the divided sample set were visually expressed. And using the above three methods, the updated samples set were selected for updating the initial discrimination models, and the spectra features of the updated samples set selected based on different methods were demonstrated. The results show that the models constructed using the LLR–SIE method after sample partitioning obtained an accuracy of 92.7% and 91% in predicting the two batches of samples, and reached a double high in both recall and precision, with both metrics above 0.9. This improves the prediction accuracy by at least 4% relative to the accuracy obtained from the models constructed after dividing the samples set using KS. The prediction accuracy of the models constructed after dividing the sample set using the SPXY method was 91% and 86.5%, and the recall and precision were low. In terms of model updating, when different numbers of updated samples are used for model updating, the accuracy obtained from updating the discrimination models by selecting updated samples using the LLR–SIE method is higher than that obtained by updating the model using the KS method and the SPXY method. After 150 updated samples were selected, the use of the LLR–SIE method resulted in the updated model prediction accuracy of 90% on the F-Measure metric. The results proved that the prediction accuracy of the model built after sample partitioning using the LLR–SIE method is better than that built after partitioning by the two traditional methods. In addition, with the same number of updated samples selected, the updated samples selected using the LLR–SIE method are more effective compared to the traditional methods, resulting in higher comprehensive evaluation indexes of the updated model. The LLR–SIE method combined with the NIR technique can provide a new solution to the problem of model construction and updates.