In this article, a new method for predominant instrument identification in polyphonic music using deep hybrid neural networks (DHNNs) is presented. The proposed deep hybrid neural network architecture uses convolutional neural networks (CNNs) for feature extraction and traditional machine learning (ML) algorithms for classification. To enhance the recognition performance further, components like batch normalization, max pooling, and global average pooling layers have also been used. Additionally, experiments are run to evaluate the performance of several deep hybrid neural network-based classifiers. In experiments, the IRMAS dataset's input audio is utilized to extract the Mel-spectrogram, which is used as a single input feature to the proposed DHNNs. The precision, recall, and F1 measures are averaged on a micro and macro scale for the evaluation of proposed DHNN model performance. By finding the optimal values for various hyperparameters through a development set with an optimal machine learning classifier, the proposed model can attain F1 measures of 0.642 and 0.564 for micro and macro, respectively, which are 3.71 and 9.94% higher than those attained by Han et al.’s CNN architecture (Han et al. in IEEE/ACM Trans. Audio. Speech Lang. Process. 25:208–221, 2016). We hope the proposed deep hybrid neural network architecture can have more prospects in various music information retrieval (MIR) applications in the future.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predominant Musical Instrument Identification Using Deep Hybrid Neural Networks

  • Sukanta Kumar Dash,
  • S. S. Solanki,
  • Soubhik Chakraborty

摘要

In this article, a new method for predominant instrument identification in polyphonic music using deep hybrid neural networks (DHNNs) is presented. The proposed deep hybrid neural network architecture uses convolutional neural networks (CNNs) for feature extraction and traditional machine learning (ML) algorithms for classification. To enhance the recognition performance further, components like batch normalization, max pooling, and global average pooling layers have also been used. Additionally, experiments are run to evaluate the performance of several deep hybrid neural network-based classifiers. In experiments, the IRMAS dataset's input audio is utilized to extract the Mel-spectrogram, which is used as a single input feature to the proposed DHNNs. The precision, recall, and F1 measures are averaged on a micro and macro scale for the evaluation of proposed DHNN model performance. By finding the optimal values for various hyperparameters through a development set with an optimal machine learning classifier, the proposed model can attain F1 measures of 0.642 and 0.564 for micro and macro, respectively, which are 3.71 and 9.94% higher than those attained by Han et al.’s CNN architecture (Han et al. in IEEE/ACM Trans. Audio. Speech Lang. Process. 25:208–221, 2016). We hope the proposed deep hybrid neural network architecture can have more prospects in various music information retrieval (MIR) applications in the future.