错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HiFi-WaveGAN: Generative Adversarial Network with Auxiliary Spectrogram-Phase Loss for High-Fidelity Singing Voice Generation

  • Chunhui Wang,
  • Chang Zeng,
  • Jun Chen,
  • Ouyang Xue

摘要

Entertainment-oriented singing voice synthesis (SVS) requires a vocoder to generate high-fidelity (e.g. 48 kHz) audio. However, most text-to-speech (TTS) vocoders cannot reconstruct the waveform well in this scenario. In this paper, we propose HiFi-WaveGAN to synthesize the 48 kHz high-quality singing voices in real-time. Specifically, it consists of an Extended WaveNet that served as a generator, a multi-period discriminator proposed in HiFiGAN, and a multi-resolution spectrogram discriminator borrowed from UnivNet. To better reconstruct the high-frequency part from the full-band mel-spectrogram, we incorporate a pulse extractor to generate the constraint for the synthesized waveform. Additionally, an auxiliary spectrogram-phase loss is utilized to approximate the real distribution further. The experimental results (Demo page: https://wavelandspeech.github.io/hifi-wavegan/ ) show that our proposed HiFi-WaveGAN obtains 4.23 in the mean opinion score (MOS) metric for the 48 kHz SVS task, significantly outperforming other neural vocoders.