ChoirDiff: Choir Vocal Synthesis via Diffusion Model
摘要
Singing voice synthesis (SVS) is to generate high quality singing audio given from lyrics and/or musical scores. Recent development in denoising diffusion probabilistic models leads to a new learning-based SVS model, DiffSinger, which can synthesize highly realistic and expressive voices. However, there is inadequate study on choral singing voice synthesis based on deep learning technology. Choral singing poses several unique challenges such as larger pitch fluctuations and increased time variations. Furthermore, there exists insufficient publicly available choral datasets for training diverse aspects of choral singing. In this paper, we built a small Latin choral dataset first and then explore the use of diffusion model like DiffSinger for Latin and Spanish choral voice synthesis (ChoirDiff) while compared to an English pop dataset with manually annotated phoneme timings. We first recorded 6 Latin songs and downloaded publicly available Cantoria Spanish choral dataset. Then we obtained their lyrics online and aligned audio to their respective lyrics phonetically and extract phoneme timings via Montreal Forced Aligner (MFA). Next three different DiffSinger models were trained with identical configurations on the extracted phoneme timings and audio to generate synthesized vocals. Mean Squared Error (MSE) was calculated to evaluate the accuracies of predictions. The models trained on choral datasets can generate realistic choral singing and achieve MSEs 0.241 for Latin, 0.143 for Spanish, and 0.120 for English, respectively. Experimental results of this paper show that DiffSinger can synthesize convincing choral vocals, and MFA is less effective in aligning complex phoneme structures present in choral vocals.