ChildTinyTalks (CTT): A Benchmark Dataset and Baseline for Expressive Child Speech Synthesis
摘要
Designing expressive speech synthesis for child voice remains an unresolved problem. One of the major dilemmas faced by child TTS systems and child speech synthesis is the scarcity of datasets to train opaque data-hungry DNN-based models. Only a few datasets were proposed for the purpose of building child conversational AI agents, and many of them come with challenges such as noisy data and indiscernible speech. With this in mind, we introduce the ChildTinyTalks (CTT) dataset, comprising 2 h of speech collected from 25 kids in grades ranging from third to fourth grade, who are telling stories and sharing their experiences. The new dataset containing 1200 audio samples has been transcribed at the word level, comprising 4 classes of voice expressions. To verify the effectiveness of CTT in real-world situations, AutoVocoder models were trained and synthesized samples were generated. The models were trained on both the LJSpeech large scale dataset and our CTT dataset. Initial experimental results indicate that the CTT dataset can steadily give comparable results with acoustic model trained on a large-scale dataset with a size of less than 10% of the large dataset.