This paper presents a solution for generating corpora of simulated Polish speech recordings in complex acoustic environments. The proposed method introduces a layer of unpredictable sound events, in addition to the acoustic scene noise and reverberation, making the solution unique. Each sound layer is stored in separate files, allowing users to mute specific layers selectively via phase cancellation. We applied this technique for data augmentation and trained two speech enhancement models. Experimental results show that the models trained with our data augmentation strategy effectively generalize across various background noise complexities. Moreover, we highlight the crucial role of integrating speech enhancement methods within the speech separation pipeline in conditions characterized by diverse background noises. Our publicly available code allows researchers to create their corpora tailored to the Polish language and train speech enhancement or separation models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Solution for Developing Corpora for Polish Speech Enhancement in Complex Acoustic Environments

  • Mariusz Kleć,
  • Krzysztof Szklanny,
  • Alicja Wieczorkowska

摘要

This paper presents a solution for generating corpora of simulated Polish speech recordings in complex acoustic environments. The proposed method introduces a layer of unpredictable sound events, in addition to the acoustic scene noise and reverberation, making the solution unique. Each sound layer is stored in separate files, allowing users to mute specific layers selectively via phase cancellation. We applied this technique for data augmentation and trained two speech enhancement models. Experimental results show that the models trained with our data augmentation strategy effectively generalize across various background noise complexities. Moreover, we highlight the crucial role of integrating speech enhancement methods within the speech separation pipeline in conditions characterized by diverse background noises. Our publicly available code allows researchers to create their corpora tailored to the Polish language and train speech enhancement or separation models.