Enhancing Hindi–English Direct Speech-to-Speech Translation with Clustering-Aided Cross-Contrastive Self-Supervised Speech Representation Learning
摘要
Direct speech-to-speech translation (S2ST) is an important tool for bridging communication gaps. Direct S2ST translates speech from one language to another without relying on intermediate text, making it particularly useful for languages primarily spoken rather than written. However, the performance of Direct S2ST models on low-resource languages remains limited due to the scarcity or complete absence of parallel speech data required for training. Pretraining and finetuning are widely used techniques to leverage unsupervised speech data to improve model performance. In this work, we employ a cluster-aided, cross-contrastive self-supervised learning (SSL)-based speech representation model as the pre-trained encoder, combined with a multilingual BART (mBART) decoder. The resulting finetuned model outperforms a baseline that uses a contrastive-loss-based SSL model as the encoder. The proposed models improve the BLEU score by 4.14% for Hindi