错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-resolution Analysis Based Audio Source Separation with Optimized U-NET Model

  • Baishakhi Dutta,
  • Chandrakant Gaikwad

摘要

Audio source separation (ASS) is a technique well-known for extracting the individual signal of underlying sound sources from the signal mixture, which is a hectic challenge due to the presence of multi-channel signals. Though there are numerous features associated with the existing separation models, time-domain signals and phase information are not considered in the conventional techniques. On the other hand, separation models working with DNN are not so efficient with sampling and the separation accuracy of the signal. To overcome these drawbacks, the optimized U-Net model with the Hybrid Wolf Optimization (HWO-UNet) is proposed in this research, which aids in separating the audio source with better accuracy. Utilizing the Hybrid optimized U-Net model with skip connection integrates the spectrogram features and the statistical features effectively for eliminating the loss of signal features and assists in the effective separation of the audio signal as well as the noise signal. The proposed method adopts the U-NET-based architecture with optimized intermediate spectrogram transformation blocks utilizing the adaptive tuning behavior of Deep CNN. Further, the U-Net acts as an encoder–decoder architecture that effectively works as an oversampling technique that mitigates the class imbalance problem and the anti-aliasing technique enhances the efficiency. The performance of the HWO-UNet ASS method in terms of Root Mean Square Error (RMSE), and Mean Square Error (MSE) is 1.242, and 0.912, respectively for the MUSDB18 whereas the maximal Signal-to-Noise Ratio (SNR) is reported as 16.51 dB for UrbanSound8k dataset.