Wheat diseases are one of the main factors that affect healthy growth, yield, and quality, leading to an annual global reduction in grain production of approximately 14%, with a growing trend. In this study, we propose the SAPS-ViM network, a novel method that integrates a bidirectional state-space model with an advanced convolutional neural network (CNN) architecture. SAPS-ViM utilizes Multi-Scale Depthwise Convolutions (MDWC) and Multi-Step Adaptive Gated Aggregation (MAGA), which are fused into a Spatial Aggregation Block (SAB) and integrated into the Vision Mamba framework. This synergistic combination improves the capability to learn features of the network by efficiently capturing contextual relationships and spatial information. Experimental results show that SAPS-ViM achieves a favorable balance between model complexity and performance through the effective coordination of multi-step depthwise convolution, adaptive gating mechanisms, and the bidirectional state-space model. Evaluation of the WPDD and LWDCD-Pro datasets demonstrates that our method achieves a classification accuracy improvement of 3.06% and 2.43%, respectively, compared to Vision Mamba, with an average accuracy increase of 4.13% over other mainstream networks of similar parameter scales. SAPS-ViM sets a new industry benchmark with superior classification accuracy, significantly exceeding established methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SAPS-ViM: Spatial Aggregation Prefix Synergistic Vision Mamba for Wheat Diseases Classification

  • Siyuan Qin,
  • Jinsong Wu

摘要

Wheat diseases are one of the main factors that affect healthy growth, yield, and quality, leading to an annual global reduction in grain production of approximately 14%, with a growing trend. In this study, we propose the SAPS-ViM network, a novel method that integrates a bidirectional state-space model with an advanced convolutional neural network (CNN) architecture. SAPS-ViM utilizes Multi-Scale Depthwise Convolutions (MDWC) and Multi-Step Adaptive Gated Aggregation (MAGA), which are fused into a Spatial Aggregation Block (SAB) and integrated into the Vision Mamba framework. This synergistic combination improves the capability to learn features of the network by efficiently capturing contextual relationships and spatial information. Experimental results show that SAPS-ViM achieves a favorable balance between model complexity and performance through the effective coordination of multi-step depthwise convolution, adaptive gating mechanisms, and the bidirectional state-space model. Evaluation of the WPDD and LWDCD-Pro datasets demonstrates that our method achieves a classification accuracy improvement of 3.06% and 2.43%, respectively, compared to Vision Mamba, with an average accuracy increase of 4.13% over other mainstream networks of similar parameter scales. SAPS-ViM sets a new industry benchmark with superior classification accuracy, significantly exceeding established methods.