<p>The task of personalized speech synthesis aims to generate speech that mimics the voice characteristics of a specific speaker. Recent advancements in large speech models, such as VALL-E, have achieved timbre cloning using a 3-second reference audio. However, current methods are limited by the reverberation and background noise in the reference audio, which can lead to unwanted information leakage into the timbre and make disentangling speaker characteristics challenging. This paper proposes a method—personalized Lao synthesis via a disentangled neural encoder-decoder language model. We present an adversarial speaker classifier and employ a mutual information minimization approach using the variational contrastive log-ratio upper bound to ensure that only the desired information features are retained during training. This approach enables more adaptive personalized Lao speech synthesis. After conducting experiments on approximately 100&#xa0;h of Lao speech data, the personalized audio synthesized based on this method achieved a MOS score of 4.02, an improvement of 0.18 compared to the baseline model VALL-E, thereby enhancing the ability to model timbre characteristics of unseen speakers.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Personalized Lao language synthesis via disentangled neural codec language model

  • Cunli Mao,
  • Tian Tian,
  • Linqin Wang,
  • Zhengtao Yu,
  • Shengxiang Gao,
  • Ling Dong

摘要

The task of personalized speech synthesis aims to generate speech that mimics the voice characteristics of a specific speaker. Recent advancements in large speech models, such as VALL-E, have achieved timbre cloning using a 3-second reference audio. However, current methods are limited by the reverberation and background noise in the reference audio, which can lead to unwanted information leakage into the timbre and make disentangling speaker characteristics challenging. This paper proposes a method—personalized Lao synthesis via a disentangled neural encoder-decoder language model. We present an adversarial speaker classifier and employ a mutual information minimization approach using the variational contrastive log-ratio upper bound to ensure that only the desired information features are retained during training. This approach enables more adaptive personalized Lao speech synthesis. After conducting experiments on approximately 100 h of Lao speech data, the personalized audio synthesized based on this method achieved a MOS score of 4.02, an improvement of 0.18 compared to the baseline model VALL-E, thereby enhancing the ability to model timbre characteristics of unseen speakers.