Analyzing cross-language similarities to enhance low-resource text-to-speech via transfer learning, case study: the Moroccan Berber Amazigh
摘要
Standard Moroccan Amazigh (Moroccan Berber) lacks any text-to-speech system despite its rich phonemic inventory. High-quality parallel data are scarce and costly for this low-resource language. We introduced a cross-lingual transfer learning approach that leverages phonemic similarity: after ranking related languages via Pearson correlation of shared phonemes, we initialized phoneme embeddings from pretrained models fine-tuned on those auxiliaries. Our TTS comprises a spectrogram predictor and neural vocoder, trained on a small dataset of Amazigh. We evaluated four strategies:Ar-Am SpeechT5 pretrained on Arabic, then fine-tuned on Amazigh; Pr-Am SpeechT5 pretrained on Persian, then fine-tuned on Amazigh; Pr-Ar-Am SpeechT5 pretrained on Persian, subsequently fine-tuned on Arabic and finally on Amazigh and En-Am SpeechT5 pretrained on English, then fine-tuned on Amazigh. Sequential cross-language transfer outperformed all other models achieving a MOS score of 3.92 the best among all evaluated approaches, demonstrating its value for low-resource TTS.