Despite the recent success of transformer architectures in single-channel speech enhancement, the existing models are computationally expensive due to many parameters and, thus, are unsuitable for TinyML applications. This chapter presents a fast framework using an Audio Spectrogram Transformer (AST) with attention layer-free encoders. The proposed model has used a discrete fractional Fourier transform to mix the input tokens, which are usually embedded vectors of spectrogram patches in speech processing. The model has been trained to estimate the a priori SNR from input noisy speech for extracting clean speech from background noise. This way, our method integrates the deep learning and signal processing tools based on two different aspects with two crucial goals. Experimental results on standard speech enhancement datasets demonstrate the effectiveness of the proposed framework, outperforming existing methods in terms of various objective metrics. The experiments reveal that, on average, the proposed approach yields 38.46% enhancement in PESQ (Perceptual Evaluation of Speech Quality) and an 11.76% improvement in STOI (Short-Time Objective Intelligibility) compared to existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Fast and Efficient Speech Enhancement Framework Using Audio Spectrogram Transformer with Attention Layer-Free Encoders

  • Suman Samui,
  • Swagata Mandal

摘要

Despite the recent success of transformer architectures in single-channel speech enhancement, the existing models are computationally expensive due to many parameters and, thus, are unsuitable for TinyML applications. This chapter presents a fast framework using an Audio Spectrogram Transformer (AST) with attention layer-free encoders. The proposed model has used a discrete fractional Fourier transform to mix the input tokens, which are usually embedded vectors of spectrogram patches in speech processing. The model has been trained to estimate the a priori SNR from input noisy speech for extracting clean speech from background noise. This way, our method integrates the deep learning and signal processing tools based on two different aspects with two crucial goals. Experimental results on standard speech enhancement datasets demonstrate the effectiveness of the proposed framework, outperforming existing methods in terms of various objective metrics. The experiments reveal that, on average, the proposed approach yields 38.46% enhancement in PESQ (Perceptual Evaluation of Speech Quality) and an 11.76% improvement in STOI (Short-Time Objective Intelligibility) compared to existing methods.