Bnvits: a voice cloning approach for single speaker text-to-speech
摘要
Although significant progress has been made in voice cloning and text-to-speech (TTS) models, especially in generating natural-sounding speech, low-resource languages such as Bangla (Bn) and other languages have remained nearly unexplored. Despite recent advancements, TTS systems for the Bangla language have still been challenged by the intricate phonology and morphology. Furthermore, no previous work had been done on voice cloning for Bangla. To address the research gap, a voice cloning method has been proposed that utilizes the limited amount of speech data available to build a TTS system for Bangla. Additionally, PYBANGLA, a text normalization tool created especially for Bangla language processing, has been introduced. Voice cloning has been achieved by refining the top-performing TTS models using just a few target speaker samples. Both subjective and objective evaluation metrics have been conducted to assess the system, and the results show that our BnVITS model has performed better than the earlier Bangla TTS model. This approach has opened up new opportunities for individualized voice technology by paving the road for more efficient Bangla TTS approaches in terms of speech data.