<p>Voice conversion (VC) requires precise preservation of source speech content while effectively capturing target speaker characteristics. Current approaches predominantly focus on extracting granular acoustic representations, yet often neglect the crucial integration of heterogeneous feature types, resulting in an inherent performance trade-off between content fidelity and speaker similarity. To address this limitation, we present IAFF-VC, a novel framework for non-parallel any-to-any voice conversion that synergistically combines encoder–decoder architecture with multi-scale feature fusion. Our proposed EMAFF module introduces a multi-branch channel grouping mechanism that strategically reorganizes spatial-semantic features across distinct subspaces, enabling optimal fusion of complementary speech attributes. Comprehensive evaluations on the VCTK benchmark demonstrate IAFF-VC’s superior performance in one-shot conversion scenarios, achieving state-of-the-art results with 9.75% CER and 91.76% speaker similarity score, while maintaining 4.21 mean opinion score for speech naturalness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

IAFF-VC: Any-to-Any Voice Conversion Using Attentional Feature Fusion

  • Kai Guo,
  • Yang Xu,
  • Sicong Zhang

摘要

Voice conversion (VC) requires precise preservation of source speech content while effectively capturing target speaker characteristics. Current approaches predominantly focus on extracting granular acoustic representations, yet often neglect the crucial integration of heterogeneous feature types, resulting in an inherent performance trade-off between content fidelity and speaker similarity. To address this limitation, we present IAFF-VC, a novel framework for non-parallel any-to-any voice conversion that synergistically combines encoder–decoder architecture with multi-scale feature fusion. Our proposed EMAFF module introduces a multi-branch channel grouping mechanism that strategically reorganizes spatial-semantic features across distinct subspaces, enabling optimal fusion of complementary speech attributes. Comprehensive evaluations on the VCTK benchmark demonstrate IAFF-VC’s superior performance in one-shot conversion scenarios, achieving state-of-the-art results with 9.75% CER and 91.76% speaker similarity score, while maintaining 4.21 mean opinion score for speech naturalness.