错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Supervised single-channel dual domains speech enhancement technique using bidirectional long short-term memory

  • Md. Shakhawat Hosen,
  • Samiul Basir,
  • Md. Farukuzzaman Khan,
  • A.O.M Asaduzzaman,
  • Md. Mojahidul Islam,
  • Md Shohidul Islam

摘要

This article introduces a brand new supervised method to make speech sound better and clearer by getting rid of background noise and mess-ups. It brings together two kinds of signal processes: the dual-tree complex wavelet transform (DTCWT) along with the short-time Fourier transform (STFT), in addition to a bidirectional long short-term memory (BLSTM) network. By taking the features from both time-frequency and wavelet domains, this way deals with issues of the discrete wavelet transform (DWT), like dealing with big changes and getting the direction right. The DTCWT makes subband signals that then change into a complex spectrogram using STFT. The size of this spectrogram is put into a bunch of layers in the BLSTM network, which finds speech bits well by knowing what came before and what’s coming next in data sequences. Soft mask parts are used to tweak this procedure. Matrix multiplication is implemented between the noisy complex spectrogram and the predicted mask of the model to produce the enhanced matrix. The final speech parts are made using inverse STFT and inverse DTCWT with the original parameters. How well this way works is tested against different kinds of noise, like stationary, quasi-stationary, and non-stationary, using both the IEEE Corpus and Aurora-2 databases. Results prove it gets a PESQ score of 3.34 and an STOI score of 0.918 at 0 signal-to-noise ratio (SNR) and also performs better for others. Other ways to gauge how good it is, like HASQI and HASPI, also do better than other methods, showing off better voice quality and understandability in different noisy spots.