Automatic musical instrument recognition has received increased attention in areas of music information retrieval (MIR) and user interface development. In the realm of music information retrieval (MIR), identifying musical instruments in polyphonic music is extremely difficult because real-life music involves a great degree of variability in terms of audio quality, timbre, and playing style. The presence of several instruments together in polyphonic music further escalates the problem. In this proposed work, we addressed an automatic predominant instrument recognition (PIR) in polyphonic music using CNN-ML framework-based network architecture through Mel-spectrogram. Our proposed CNN-ML network uses convolutional neural networks (CNNs) for feature extraction and machine learning (ML) algorithms for instrument classification. Experiments are carried out to evaluate the performance of different proposed CNN-ML networks by fusing different ML classifiers with the cutting-edge Han’s CNN model. In experiments, the Mel-spectrogram feature of input audio is extracted using the IRMAS dataset. By averaging the precision, recall, and F1 metrics on a micro and macro level, the proposed network's performance is evaluated. For an optimal ML classifier combined with the cutting-edge Han’s CNN model, our proposed CNN-ML network architecture can reach F1 metrics on micro and macro levels as 0.626 and 0.526, respectively, which outperforms those obtained by the cutting-edge Han's CNN model by 1.13% and 2.53%, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CNN-ML Framework-Based Predominant Musical Instrument Recognition Using Mel-Spectrogram

  • Sukanta Kumar Dash,
  • S. S. Solanki,
  • Soubhik Chakraborty

摘要

Automatic musical instrument recognition has received increased attention in areas of music information retrieval (MIR) and user interface development. In the realm of music information retrieval (MIR), identifying musical instruments in polyphonic music is extremely difficult because real-life music involves a great degree of variability in terms of audio quality, timbre, and playing style. The presence of several instruments together in polyphonic music further escalates the problem. In this proposed work, we addressed an automatic predominant instrument recognition (PIR) in polyphonic music using CNN-ML framework-based network architecture through Mel-spectrogram. Our proposed CNN-ML network uses convolutional neural networks (CNNs) for feature extraction and machine learning (ML) algorithms for instrument classification. Experiments are carried out to evaluate the performance of different proposed CNN-ML networks by fusing different ML classifiers with the cutting-edge Han’s CNN model. In experiments, the Mel-spectrogram feature of input audio is extracted using the IRMAS dataset. By averaging the precision, recall, and F1 metrics on a micro and macro level, the proposed network's performance is evaluated. For an optimal ML classifier combined with the cutting-edge Han’s CNN model, our proposed CNN-ML network architecture can reach F1 metrics on micro and macro levels as 0.626 and 0.526, respectively, which outperforms those obtained by the cutting-edge Han's CNN model by 1.13% and 2.53%, respectively.