VCIVS: Integrating Voice Cloning Technology to Improve Vietnamese Storytelling
摘要
This research explores the landscape of Vietnamese storytelling through the integration of the cutting-edge voice cloning technology. Rigorous experiments with Multispeaker Fastspeech2 and VITS, utilizing evaluation metrics such as DNSMOS and GE2E, lead to the selection of Multispeaker Fastspeech2 for its superior synthesis capabilities, with an impressive DNSMOS score of 4.3, ensuring an authentic and emotionally resonant storytelling experience within the framework of this study. Combined with the pre-trained HiFi-GAN model, a swift and lightweight mel-spectrogram decoder, along with the pre-trained ResNet34 model for robust speaker feature encoding. These integrations into a storytelling system, score of 4.2 in MOS in the overall satisfaction, contribute to an innovative and culturally enriched storytelling experience for both children and parents. A preview of the voice cloning quality can be accessed through the link https://dghthuong94.github.io/VCIVS .