Rectangular kernels for information-dense domains in environmental sound classification
摘要
The mainstream approaches for environmental sound classification (ESC) typically employ audio spectrograms as input and leverage techniques from image processing. However, a key distinction exists between audio spectrograms and images: the former are generally rectangular, whereas the latter are commonly square or nearly square. This rectangular form stems from the inherent structure of audio signals, which span both time and frequency domains. Notably, the frequency dimension of spectrograms is often more compact and information-rich than the time dimension. In light of this, we propose to replace the conventional square convolution kernel–which treats both dimensions equally–with rectangular kernels combined with dilated convolutions. This design prioritises the more informative frequency axis. We further introduce a neural network model tailored for ESC, enhanced by self-distilled soft labels to enrich the input information, and a reconstructed loss function that boosts both accuracy and robustness. Experimental results on UrbanSound8K, ESC-10, and ESC-50 datasets achieve accuracies of 98.62%, 95.50%, and 89.3%, respectively, matching or surpassing state-of-the-art performance while maintaining a lower parameter count–demonstrating the efficiency of our approach.