错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Voice Cloning with RVC and Edge TTS: Enhancing Accuracy Through Algorithmic Integration

  • Rakshitha,
  • Rashmitha Shettigar,
  • M. S. Krishna,
  • Kushi S. Bhimani,
  • D. N. Disha,
  • Sudesh Rao

摘要

Voice cloning has emerged as a transformative technology with applications ranging from entertainment and personalized digital assistants to accessibility solutions for individuals with speech impairments. Despite its growing adoption, existing systems often struggle to achieve high levels of naturalness, fidelity, and adaptability across diverse languages and accents. This research addresses these limitations by strategically integrating Retrieval-based Voice Conversion (RVC V2) with Microsoft’s Edge text-to-speech (TTS). RVC provides robustness and adaptability in capturing voice characteristics, while Edge TTS contributes wide multilingual and accent support, low-latency performance, and expressive intonation. Together, this hybrid framework enhances synthesis accuracy, realism, and flexibility. To further optimize performance, advanced algorithms such as ContentVec, VITS, HIFIGAN, UVR5, and RMVPE are employed, enabling improvements in tone modulation, clarity, and overall fidelity. Key contributions of this work include the novel integration of TTS within a voice cloning pipeline, the incorporation of multiple complementary algorithms, and the exploration of individualized training for tailored voice models. The system was evaluated through subjective listening sessions, where participants consistently reported substantial improvements in naturalness and similarity compared to outputs from standalone RVC models. These results confirm the effectiveness of combining TTS and voice cloning, with the hybrid approach producing voices nearly indistinguishable from human speech. Beyond technical improvements, the research underscores the broader significance of voice cloning advancements for accessibility, human–computer interaction, and content creation. By demonstrating a scalable, accurate, and flexible framework, this study contributes to the next generation of synthetic voice systems, capable of enriching communication and delivering lifelike digital experiences.