An audio generation model based on empirical mode decomposition and generative adversarial networks for enhancing voice quality and diversity
摘要
Since the advent of generative adversarial networks (GANs), various studies have been conducted to generate new samples by imitating biological samples. These generated samples find applications in various fields, including data augmentation and voice conversion. However, the quality and diversity of the generated audio can still be improved. In this study, we propose EMDGAN, an audio generation method based on improved complete ensemble empirical mode decomposition (ICEEMD) and GANs, which imitates the intrinsic mode functions of speech and eventually generates speech with better quality and more diversity. Furthermore, this paper proposes a two-stage filtering process that effectively selects high-quality speech from a large volume of generated speech samples. To evaluate the performance, we employ objective measures including the inception score and Fréchet inception distances, alongside a subjective test to assess the clarity and naturalness of the generated samples. Additionally, we utilize the t-distributed random neighborhood embedding visualization tool to compare the diversity of the generated samples. Experimental results demonstrate that EMDGAN generates speech samples with greater diversity and improved quality compared to WaveGAN. In the application of data augmentation from a small dataset for speech recognition, the employment of EMDGAN proves effective in enhancing word recognition rates.