<p>A variety of single-channel speech enhancement techniques, including CRN (Convolutional Recurrent Networks), TCNN (Temporal Convolutional Neural Network), WAVUNET (Waveform U-Net), ICCRN (Iterative Cross-channel Recurrent Network), SMWETCW (Sparse Multi-Wavelet Transform Coupled with Waveform), IUINNB (Improved Unsupervised Incremental Neural Network for Big Data), TDSTN (Temporal Dynamic Spatial Transformer Network), and MSP (Multiscale Perceptron), demonstrate specific limitations. TCNN and WAVUNET exhibit aliasing issues stemming from exponential dilation rates, which limit their ability to preserve critical speech attributes. SMWETCW introduces latency by explicit windowing, while IUINNB increases model complexity by incorporating uncertainty modeling. We present a system for improving speech that merges MS-SENet (Multi-Scale Squeeze-and-Excitation Network) with a Lightweight Multi-Axial Transformer, D3 Net, and U-Net, called MS-SENet (LMATD3MUNet). MS-SNet, or Multi-Scale Squeeze Network, is a framework that employs a multi-scale approach to effectively compress and process data. The Embedded U-Netset is a decoder-encoder framework designed to enhance data flow efficiency. In LMATD3MUNet, the D3Net block incorporates a Multi-Scale Feature Extractor (MSFE) to effectively accumulate comprehensive contextual information. We integrated a D3Net block with an MSFE block in MS-SENetLMATD3MUNet to obtain comprehensive pertinent data. This enables the extensive utilization of both global and local features, substantially enhancing speech reconstruction abilities. Employing a novel technique termed multi-dilated convolution, characterized by adjustable dilation values for each layer. D3Net simultaneously emulates many resolutions. D3Net improves the enhancement of free space and the mathematical representation of multi-resolution data within a single convolutional layer. The MS-SENet is designed to augment the representational capability of convolutional neural networks (CNNs) by including a wider range of inputs across several resolutions. The LMAT module can enhance novel inter-channel interactions. Despite the decrease in dimensionality, the choice of an adaptable kernel size during module testing significantly improved network speed. We use Squeeze-and-Excitation (SE) modules to gather information from different scales, using skip connections and Spatial Dropout (SD) layers to reduce overfitting and improve the model's complexity. The methodology attains performance that surpasses the prior state-of-the-art technology. We integrated D3Net, LMAT, and MS-SENet into the proposed model to enhance feature extraction and contextual integration at the utterance level. The conducted experiments corroborate the proposed method. MS-SENetLMATD3MUNet outperforms benchmark models in PESQ and STOI&#xa0;measures.The proposed method improves 82.17% STOI (%) of values on average. Spanish language is giving the height value, when compared with other languages and the proposed method improves 2.81 PESQ of values on average.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight Multi-axial Transformer with MS-SENet and D3Net for Single Channel Speech Enhancement

  • Silpa Peethala,
  • V. Sunnydayal

摘要

A variety of single-channel speech enhancement techniques, including CRN (Convolutional Recurrent Networks), TCNN (Temporal Convolutional Neural Network), WAVUNET (Waveform U-Net), ICCRN (Iterative Cross-channel Recurrent Network), SMWETCW (Sparse Multi-Wavelet Transform Coupled with Waveform), IUINNB (Improved Unsupervised Incremental Neural Network for Big Data), TDSTN (Temporal Dynamic Spatial Transformer Network), and MSP (Multiscale Perceptron), demonstrate specific limitations. TCNN and WAVUNET exhibit aliasing issues stemming from exponential dilation rates, which limit their ability to preserve critical speech attributes. SMWETCW introduces latency by explicit windowing, while IUINNB increases model complexity by incorporating uncertainty modeling. We present a system for improving speech that merges MS-SENet (Multi-Scale Squeeze-and-Excitation Network) with a Lightweight Multi-Axial Transformer, D3 Net, and U-Net, called MS-SENet (LMATD3MUNet). MS-SNet, or Multi-Scale Squeeze Network, is a framework that employs a multi-scale approach to effectively compress and process data. The Embedded U-Netset is a decoder-encoder framework designed to enhance data flow efficiency. In LMATD3MUNet, the D3Net block incorporates a Multi-Scale Feature Extractor (MSFE) to effectively accumulate comprehensive contextual information. We integrated a D3Net block with an MSFE block in MS-SENetLMATD3MUNet to obtain comprehensive pertinent data. This enables the extensive utilization of both global and local features, substantially enhancing speech reconstruction abilities. Employing a novel technique termed multi-dilated convolution, characterized by adjustable dilation values for each layer. D3Net simultaneously emulates many resolutions. D3Net improves the enhancement of free space and the mathematical representation of multi-resolution data within a single convolutional layer. The MS-SENet is designed to augment the representational capability of convolutional neural networks (CNNs) by including a wider range of inputs across several resolutions. The LMAT module can enhance novel inter-channel interactions. Despite the decrease in dimensionality, the choice of an adaptable kernel size during module testing significantly improved network speed. We use Squeeze-and-Excitation (SE) modules to gather information from different scales, using skip connections and Spatial Dropout (SD) layers to reduce overfitting and improve the model's complexity. The methodology attains performance that surpasses the prior state-of-the-art technology. We integrated D3Net, LMAT, and MS-SENet into the proposed model to enhance feature extraction and contextual integration at the utterance level. The conducted experiments corroborate the proposed method. MS-SENetLMATD3MUNet outperforms benchmark models in PESQ and STOI measures.The proposed method improves 82.17% STOI (%) of values on average. Spanish language is giving the height value, when compared with other languages and the proposed method improves 2.81 PESQ of values on average.