<p>The automated analysis of bird vocalizations is crucial for applications in ecology, conservation, and vocal behavioural studies. Recent advancements in deep learning have improved classification accuracy, but existing models often struggle with small and imbalanced datasets. This paper presents a novel multi-label bird audio classification framework using Res2Net combined with Convolutional Block Attention Module (CBAM), Spatial Attention (SA), and Grouped Channel Feature Attention (GCA) employing a sequential aggregation (Se) strategy. The proposed model refines predictions by normalizing aggregated sigmoid probabilities and selecting target species based on the highest scores. We evaluate our approach on the xeno-canto database, consisting of 10 bird species. We tested the proposed model on 434 audio recordings with calls of two and three species. The presented framework, which is characterized by fewer parameters compared to other residual networks with attention mechanism, can be directly used for image and audio classification tasks. Experimental results demonstrate that our model achieves an F1-score of 72.20% using Mel-spectrogram features, outperforming state-of-the-art methods. The novelty of the proposed framework lies in the sequential aggregation strategy and Residual CBAM with grouped channel attention mechanism.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identifying overlapping bird species from raw field audio recordings by assembling Grouped Channel Feature Attention with multi-scale residual CBAM

  • Noumida Abdul Kareem,
  • Rajeev Rajan

摘要

The automated analysis of bird vocalizations is crucial for applications in ecology, conservation, and vocal behavioural studies. Recent advancements in deep learning have improved classification accuracy, but existing models often struggle with small and imbalanced datasets. This paper presents a novel multi-label bird audio classification framework using Res2Net combined with Convolutional Block Attention Module (CBAM), Spatial Attention (SA), and Grouped Channel Feature Attention (GCA) employing a sequential aggregation (Se) strategy. The proposed model refines predictions by normalizing aggregated sigmoid probabilities and selecting target species based on the highest scores. We evaluate our approach on the xeno-canto database, consisting of 10 bird species. We tested the proposed model on 434 audio recordings with calls of two and three species. The presented framework, which is characterized by fewer parameters compared to other residual networks with attention mechanism, can be directly used for image and audio classification tasks. Experimental results demonstrate that our model achieves an F1-score of 72.20% using Mel-spectrogram features, outperforming state-of-the-art methods. The novelty of the proposed framework lies in the sequential aggregation strategy and Residual CBAM with grouped channel attention mechanism.