<p>The detection of fraudulent pinto bean adulteration, typically driven by the introduction of low-grade specimens with the Hard-to-Cook (HTC) defect into premium batches, is a persistent challenge in the legume industry. Near-infrared (NIR) spectroscopy combined with machine learning offers a non-destructive route for cooking-quality grading, yet industrial deployment remains constrained by sensor noise, environmental variability, and the CPU-only compute budget of in-line grading hardware. In this setting, the question of which classifier to deploy is further conditioned by whether an offline hyperparameter search can be afforded or not. To address this gap, four CPU-deployable classifiers (PLS-DA, Bagging, CatBoost, and a compact Multi-Layer Perceptron) are benchmarked under two complementary configurations (out-of-the-box and GridSearchCV-tuned) against four physically motivated training-noise families (additive Gaussian, multiplicative scatter, baseline drift, and Heteroscedastic Gaussian), across six intensity levels and 100 Monte Carlo train-test splits, with all results evaluated on clean test data. The analysis yields a regime-dependent ranking: the tuned PLS-DA is the least sensitive classifier when an offline search is feasible, whereas the out-of-the-box MLP is preferable under tight deployment budgets. Under Heteroscedastic Gaussian noise the degradation pattern is broadly aligned with that of additive Gaussian noise at the same mean per-feature intensity, with the tuned PLS-DA retaining its lead but exhibiting a markedly larger Accuracy dispersion as the wavelength-dependent variance interacts with the latent decomposition. The SHAP interpretability analysis further reveals that the additional latent capacity introduced by tuning is specifically redirected towards the 1400–1500&#xa0;nm band, the first O-H overtone associated with water, which is the spectral region most directly coupled to the hydration kinetics that define the HTC defect. These findings translate into explicit, deployment-oriented guidance for classifier selection in CPU-based spectroscopic grading of pinto beans.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Training noise sensitivity analysis of default and tuned classifiers for NIR spectroscopic detection of pinto bean adulteration

  • Raziyeh Pourdarbani,
  • Sajad Sabzi,
  • Dorrin Sotoudeh,
  • Mohammadreza Ahmaditeshnizi,
  • Nadia Saadati,
  • Ruben Fernandez-Beltran,
  • Ginés García-Mateos,
  • Juan Ignacio Arribas

摘要

The detection of fraudulent pinto bean adulteration, typically driven by the introduction of low-grade specimens with the Hard-to-Cook (HTC) defect into premium batches, is a persistent challenge in the legume industry. Near-infrared (NIR) spectroscopy combined with machine learning offers a non-destructive route for cooking-quality grading, yet industrial deployment remains constrained by sensor noise, environmental variability, and the CPU-only compute budget of in-line grading hardware. In this setting, the question of which classifier to deploy is further conditioned by whether an offline hyperparameter search can be afforded or not. To address this gap, four CPU-deployable classifiers (PLS-DA, Bagging, CatBoost, and a compact Multi-Layer Perceptron) are benchmarked under two complementary configurations (out-of-the-box and GridSearchCV-tuned) against four physically motivated training-noise families (additive Gaussian, multiplicative scatter, baseline drift, and Heteroscedastic Gaussian), across six intensity levels and 100 Monte Carlo train-test splits, with all results evaluated on clean test data. The analysis yields a regime-dependent ranking: the tuned PLS-DA is the least sensitive classifier when an offline search is feasible, whereas the out-of-the-box MLP is preferable under tight deployment budgets. Under Heteroscedastic Gaussian noise the degradation pattern is broadly aligned with that of additive Gaussian noise at the same mean per-feature intensity, with the tuned PLS-DA retaining its lead but exhibiting a markedly larger Accuracy dispersion as the wavelength-dependent variance interacts with the latent decomposition. The SHAP interpretability analysis further reveals that the additional latent capacity introduced by tuning is specifically redirected towards the 1400–1500 nm band, the first O-H overtone associated with water, which is the spectral region most directly coupled to the hydration kinetics that define the HTC defect. These findings translate into explicit, deployment-oriented guidance for classifier selection in CPU-based spectroscopic grading of pinto beans.