Robust Classification of Urban Sounds in Noisy Environments: A Novel Approach Using SPWVD-MFCC and Dual-Stream Classifier
摘要
Urban sound classification is essential for effective sound monitoring and mitigation strategies, which are critical to addressing the negative impacts of noise pollution on public health. While existing methods predominantly rely on Short-Term Fourier Transform (STFT)-based features like Mel-Frequency Cepstral Coefficients (MFCC), these approaches often struggle to identify the dominant sound in noisy environments. This gap in robustness limits the practical deployment of such systems in real-world urban settings, where noise levels are unpredictable and variable. Here, we introduce Smoothed Pseudo-Wigner–Ville Distribution-based MFCC (SPWVD-MFCC), a novel feature that merges SPWVD’s high time–frequency resolution with MFCC’s human-like auditory sensitivity. We further propose a dual-stream ResNet50-CNN-LSTM architecture to classify these features. Comprehensive experiments conducted on UrbanSound8K, UrbanSoundPlus, and DCASE2016 datasets demonstrate that the proposed SPWVD-MFCC significantly improves classification accuracy in noisy conditions, with an enhancement of up to 37.2% over traditional STFT-based methods and better robustness than existing approaches. These results indicate that the proposed approach addresses a critical gap in urban sound classification by providing enhanced robustness in low-SNR environments. This advancement improves the reliability of urban noise monitoring systems and contributes to the broader goal of creating healthier urban living environments by enabling more effective noise-control strategies.