Improving Honey Adulteration Detection with Feature Selection and Resampling
摘要
Pure and reliable honey is essential. This research addresses honey adulteration detection using hyperspectral imagery, feature selection, resampling, and machine learning. Hyperspectral data from the New Zealand honey dataset, representing various honey types and adulteration levels, was analyzed using SVM, LR, and RF. The dataset was split 80–20% for training and testing, with fivefold cross-validation on the training set. Four models were evaluated: baseline, RFE-based feature selection, SMOTE for class imbalance, and SMOTE with PCA for dimensionality reduction. RF performed best, achieving a 0.996 mean accuracy, 0.997 test accuracy, and F1 scores of 0.986 (pure) and 0.998 (adulterated) for the RFE model with 75 key features. This study offers an accurate, efficient solution for honey adulteration detection, enhancing quality assessment and consumer trust.