Music Versus Speech Classification Using 3D CNN
摘要
In the Scenario of technological advancements, the accurate classification of audio signals stands as a decisive challenge related to applications like Automatic Speech Recognition (ASR), Music Information Retrieval (MIR) systems, music recognition, speaker identification, and speech enhancement. This paper proposes a novel methodology for the subtle classification of music versus speech, motivated by a three-dimensional Convolutional Neural Network (3D CNN). Unlike traditional methods, the proposed model illustrates the power of using 3D CNN on the 3D representation of Mel-Spectrograms, offering an innovative means to capture intricate temporal and spectral features inherent in audio signals. The proposed model gives good accuracy in the classification between music and speech, demonstrating its efficacy in addressing the challenges associated with audio signal classification.