错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Zero-Shot Singing Voice Conversion Based on Timbre Space Modeling and Excitation Signal Control

  • Yuan Jiang,
  • Yan-Nian Chen,
  • Li-Juan Liu,
  • Ya-Jun Hu,
  • Xin Fang,
  • Zhen-Hua Ling

摘要

In recent years, singing voice conversion technology has rapidly advanced and is capable of generating high-quality singing voices. However, challenges persist, such as pitch fluctuations and significant differences in pitch ranges between source and target singers, leading to conflicts between the accuracy and similarity of the converted pitch. Additionally, the majority of current methods require some data from the target singer to train the model, and the performance and robustness of zero-shot singing voice conversion methods have not been explored. This paper introduces new modifications within the VITS framework to achieve zero-shot singing voice conversion: 1. Timbre space modeling based on Glow; 2. Incorporating excitation signal into the decoder for waveform generation to explicitly control the pitch; 3. Proposing dual-decoder to enhance the stability of 48 kHz waveform modeling for high-quality voice generation; 4. Proposing key shift based pitch mapping strategy for conversion stage. Experimental results demonstrate that Glow-based timbre space modeling enhances the similarity in zero-shot conversion. The excitation signal module and the dual-decoder contribute to naturalness and stability. The pitch mapping strategy effectively avoids out-of-tune problems without compromising similarity to the target singer.