Personalized Lao language synthesis via disentangled neural codec language model
摘要
The task of personalized speech synthesis aims to generate speech that mimics the voice characteristics of a specific speaker. Recent advancements in large speech models, such as VALL-E, have achieved timbre cloning using a 3-second reference audio. However, current methods are limited by the reverberation and background noise in the reference audio, which can lead to unwanted information leakage into the timbre and make disentangling speaker characteristics challenging. This paper proposes a method—personalized Lao synthesis via a disentangled neural encoder-decoder language model. We present an adversarial speaker classifier and employ a mutual information minimization approach using the variational contrastive log-ratio upper bound to ensure that only the desired information features are retained during training. This approach enables more adaptive personalized Lao speech synthesis. After conducting experiments on approximately 100 h of Lao speech data, the personalized audio synthesized based on this method achieved a MOS score of 4.02, an improvement of 0.18 compared to the baseline model VALL-E, thereby enhancing the ability to model timbre characteristics of unseen speakers.