<p>Direct speech-to-speech translation (S2ST) is an important tool for bridging communication gaps. Direct S2ST translates speech from one language to another without relying on intermediate text, making it particularly useful for languages primarily spoken rather than written. However, the performance of Direct S2ST models on low-resource languages remains limited due to the scarcity or complete absence of parallel speech data required for training. Pretraining and finetuning are widely used techniques to leverage unsupervised speech data to improve model performance. In this work, we employ a cluster-aided, cross-contrastive self-supervised learning (SSL)-based speech representation model as the pre-trained encoder, combined with a multilingual BART (mBART) decoder. The resulting finetuned model outperforms a baseline that uses a contrastive-loss-based SSL model as the encoder. The proposed models improve the BLEU score by 4.14% for Hindi<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41314_2025_78_Article_IEq1.gif" Format="GIF" Height="6" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(\rightarrow \)</EquationSource> <EquationSource Format="MATHML"><math> <mo stretchy="false">→</mo> </math></EquationSource> </InlineEquation>English and 8.2% for English<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41314_2025_78_Article_IEq1.gif" Format="GIF" Height="6" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(\rightarrow \)</EquationSource> <EquationSource Format="MATHML"><math> <mo stretchy="false">→</mo> </math></EquationSource> </InlineEquation>Hindi compared to their respective baseline models. To train the model for English-to-Hindi, we trained a unit-vocoder on speech quantized using ensemble clustering instead of standard clustering. The resulting unit-vocoder outperformed the one trained on speech quantized using standard k-means for all evaluation metrics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Hindi–English Direct Speech-to-Speech Translation with Clustering-Aided Cross-Contrastive Self-Supervised Speech Representation Learning

  • Mahendra Gupta,
  • Maitreyee Dutta,
  • Chandresh Kumar Maurya

摘要

Direct speech-to-speech translation (S2ST) is an important tool for bridging communication gaps. Direct S2ST translates speech from one language to another without relying on intermediate text, making it particularly useful for languages primarily spoken rather than written. However, the performance of Direct S2ST models on low-resource languages remains limited due to the scarcity or complete absence of parallel speech data required for training. Pretraining and finetuning are widely used techniques to leverage unsupervised speech data to improve model performance. In this work, we employ a cluster-aided, cross-contrastive self-supervised learning (SSL)-based speech representation model as the pre-trained encoder, combined with a multilingual BART (mBART) decoder. The resulting finetuned model outperforms a baseline that uses a contrastive-loss-based SSL model as the encoder. The proposed models improve the BLEU score by 4.14% for Hindi \(\rightarrow \) English and 8.2% for English \(\rightarrow \) Hindi compared to their respective baseline models. To train the model for English-to-Hindi, we trained a unit-vocoder on speech quantized using ensemble clustering instead of standard clustering. The resulting unit-vocoder outperformed the one trained on speech quantized using standard k-means for all evaluation metrics.