错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ChoirDiff: Choir Vocal Synthesis via Diffusion Model

  • Xie Yun Di

摘要

Singing voice synthesis (SVS) is to generate high quality singing audio given from lyrics and/or musical scores. Recent development in denoising diffusion probabilistic models leads to a new learning-based SVS model, DiffSinger, which can synthesize highly realistic and expressive voices. However, there is inadequate study on choral singing voice synthesis based on deep learning technology. Choral singing poses several unique challenges such as larger pitch fluctuations and increased time variations. Furthermore, there exists insufficient publicly available choral datasets for training diverse aspects of choral singing. In this paper, we built a small Latin choral dataset first and then explore the use of diffusion model like DiffSinger for Latin and Spanish choral voice synthesis (ChoirDiff) while compared to an English pop dataset with manually annotated phoneme timings. We first recorded 6 Latin songs and downloaded publicly available Cantoria Spanish choral dataset. Then we obtained their lyrics online and aligned audio to their respective lyrics phonetically and extract phoneme timings via Montreal Forced Aligner (MFA). Next three different DiffSinger models were trained with identical configurations on the extracted phoneme timings and audio to generate synthesized vocals. Mean Squared Error (MSE) was calculated to evaluate the accuracies of predictions. The models trained on choral datasets can generate realistic choral singing and achieve MSEs 0.241 for Latin, 0.143 for Spanish, and 0.120 for English, respectively. Experimental results of this paper show that DiffSinger can synthesize convincing choral vocals, and MFA is less effective in aligning complex phoneme structures present in choral vocals.