A Survey of Singing Voice Synthesis
摘要
Singing voice synthesis (SVS) is a task of synthesizing singing audio according to music score and lyrics. SVS is a sub-research direction of speech synthesis. Singing voice is more complex and changeable than speaking voice in pitch, duration, and energy, so modeling singing voice is more difficult. With the rapid development of deep learning, more and more researchers in academia and industry have devoted themselves to research, and the research goal is to generate natural, expressive and human-like singing voices. This paper analyzes and compares the current challenges and progress in the research of SVS from three aspects: datasets, data augmentation strategies, and SVS models. Among them, the key point is to compare how the two-stage and fully end-to-end SVS models solve various problems in this field. Finally, we summarize the current research situation and discuss future research directions. It is expected to bring some inspiration to researchers in this field.