<p>Recent progress in music source separation has been accelerated by deep learning techniques, yet most studies have focused on Western instruments and vocals, with limited attention to traditional Chinese instruments. These instruments possess distinctive timbral characteristics and complex performance techniques, which require specialized treatment in separation tasks. This paper introduces a deep learning approach that is tuned to the frequency band to separate traditional Chinese instrument sources. By analyzing the spectral energy distributions of guzheng, dizi, pipa, and xiao, the model adopts differentiated frequency band processing strategies based on each instrument’s acoustic profile. The architecture integrates convolutional and recurrent modules with frequency attention and multi-head attention mechanisms to enhance music source separation performance. Extensive experiments with 13 band-division configurations reveal significant variations in sensitivity across instruments, with optimal frequency splits aligning closely with their spectral characteristics. The results demonstrate that the proposed method achieves high-quality music source separation while reducing computational costs through adaptive spectral processing. These findings highlight the importance of culturally informed modeling in the separation of music sources and open new directions for the preservation and analysis of traditional music. All model weights, source code, and audio demonstrations are publicly available at <a href="https://huggingface.co/NMLAB8/CISM">https://huggingface.co/NMLAB8/CISM</a> and <a href="https://huggingface.co/spaces/NMLAB8/CISM">https://huggingface.co/spaces/NMLAB8/CISM</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Chinese instrument music source separation with frequency-attentive multi-band neural networks

  • Jiaxiang Zheng,
  • Moxi Cao,
  • Chongbin Zhang

摘要

Recent progress in music source separation has been accelerated by deep learning techniques, yet most studies have focused on Western instruments and vocals, with limited attention to traditional Chinese instruments. These instruments possess distinctive timbral characteristics and complex performance techniques, which require specialized treatment in separation tasks. This paper introduces a deep learning approach that is tuned to the frequency band to separate traditional Chinese instrument sources. By analyzing the spectral energy distributions of guzheng, dizi, pipa, and xiao, the model adopts differentiated frequency band processing strategies based on each instrument’s acoustic profile. The architecture integrates convolutional and recurrent modules with frequency attention and multi-head attention mechanisms to enhance music source separation performance. Extensive experiments with 13 band-division configurations reveal significant variations in sensitivity across instruments, with optimal frequency splits aligning closely with their spectral characteristics. The results demonstrate that the proposed method achieves high-quality music source separation while reducing computational costs through adaptive spectral processing. These findings highlight the importance of culturally informed modeling in the separation of music sources and open new directions for the preservation and analysis of traditional music. All model weights, source code, and audio demonstrations are publicly available at https://huggingface.co/NMLAB8/CISM and https://huggingface.co/spaces/NMLAB8/CISM.