Cross-Lingual Speaker Transfer for Cambodian Based on Feature Disentangler and Time-Frequency Attention Adaptive Normalization
摘要
Given the scarcity of a multi-speaker corpus in Cambodian, conventional methods have shown poor performance in Cambodian speaker tranfer. On the other hand, simply using Chinese-English rich resources to expand the training data faces problems in disentangling of linguistic feature and speaker timbre feature. This paper proposes to build a cross-lingual feature disentangler and incorporate Time-Frequency Attention Adaptive Normalization (TFAAN) to effectively transform Cambodian speaker timbre into Chinese-English without altering Cambodian speech content, which enables leveraging speaker timbre features from non-parallel Chinese-English corpus as an augmentation. The experiments show that the synthesized audio achieves a MOS score of 3.81, indicating an effective disentangling and controlled transfer of speaker timbre feature.