Fusion Spectrogram for Sound Classification Using 2D Convolutional Neural Network
摘要
Due to the numerous practical applications, environmental sound classification is rapidly gaining popularity in the research field of signal processing. The research proposes an improved method for extracting specific features from audio inputs using a fusion spectrogram. A 2D convolutional neural network (2D-CNN) is used to investigate the environmental sound classification performance change, using three alternative two-dimensional time–frequency features. The evaluations classify ten classes of environmental sounds in varying degrees of acoustic background noise using spectrograms, Mel spectrograms, and combinations of these two, the proposed fusion spectrogram and displaying the exciting behavior of the classifier with the changes in input. A 2D CNN is used to extract the spatial properties, focusing on moving the spectrogram spatially over each spectral band. The proposed 2D CNN generates the outputs based on the retrieved features at the last layer of the architecture, which has ten neurons representing each class. The accuracy and effectiveness of the proposed classification are improved using a 2D CNN and batch stochastic gradient descent.