In the Scenario of technological advancements, the accurate classification of audio signals stands as a decisive challenge related to applications like Automatic Speech Recognition (ASR), Music Information Retrieval (MIR) systems, music recognition, speaker identification, and speech enhancement. This paper proposes a novel methodology for the subtle classification of music versus speech, motivated by a three-dimensional Convolutional Neural Network (3D CNN). Unlike traditional methods, the proposed model illustrates the power of using 3D CNN on the 3D representation of Mel-Spectrograms, offering an innovative means to capture intricate temporal and spectral features inherent in audio signals. The proposed model gives good accuracy in the classification between music and speech, demonstrating its efficacy in addressing the challenges associated with audio signal classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Music Versus Speech Classification Using 3D CNN

  • Pramod Hullur,
  • Dhanya Kulkarni,
  • Vinayak Jainapur,
  • Satish Chikkamath,
  • S. R. Nirmala,
  • Suneeta V. Budihal

摘要

In the Scenario of technological advancements, the accurate classification of audio signals stands as a decisive challenge related to applications like Automatic Speech Recognition (ASR), Music Information Retrieval (MIR) systems, music recognition, speaker identification, and speech enhancement. This paper proposes a novel methodology for the subtle classification of music versus speech, motivated by a three-dimensional Convolutional Neural Network (3D CNN). Unlike traditional methods, the proposed model illustrates the power of using 3D CNN on the 3D representation of Mel-Spectrograms, offering an innovative means to capture intricate temporal and spectral features inherent in audio signals. The proposed model gives good accuracy in the classification between music and speech, demonstrating its efficacy in addressing the challenges associated with audio signal classification.