Cross-Lingual Style Transfer TTS for High-Quality Machine Dubbing
摘要
Style transfer is essential for high-quality machine dubbing. While numerous approaches for style transfer in speech synthesis have been developed, cross-lingual style transfer remains a significant challenge. In this paper we introduce a novel speech synthesis method which realizes style learning across different languages and speakers. Our approach features a transformer-based architecture with a speech prompted text encoder, a duration predictor and a flow matching generative decoder. The text encoder is conditioned on the noisy source language speech, which is entered as speech prompt for style adaptation. The flow matching generative decoder produces high quality speech conditioned on the text encoder output. Empirical evaluations demonstrate that our TTS system generates speech that is objectively closer to the recordings of professional voice talents compared to a strong baseline model. Audio samples are available on our demo page ( http://lml.bas.bg/~stoyan/dubbing ).