<p>Text-To-Speech (TTS) technology has progressive substantially, yet challenges persist in achieving human-like speech quality, emotional expressiveness, and efficient multilingual synthesis. To address these issues, this research proposes the Convolutional Swin transformer-based conditional Variational Elk Herd Optimizer (CSV-EHO) system, an advanced text-augmented speech synthesis model. The system integrates text normalization with GPT-3, phoneme conversion using an adapted Grapheme-to-Phoneme (G2P) model, and Mel-spectrogram generation leveraging Swin Transformer and convolutional networks. The Elk Herd Optimizer (EHO) enhances phoneme-to-spectrogram mapping, while a dual-path WaveNet vocoder with Cosine Wavelet Transform refines clarity and naturalness. Experimental evaluations on VCTK and LJ Speech datasets demonstrate excellent performance, achieving a MOS of 4.57 ± 0.21, CMOS of + 0.45, SSIM of 0.91, phoneme accuracy of 98.2%, and WER of 3.8%, along with PESQ scores of 2.9 and 2.6. The model also attains an inference time of 0.011&#xa0;s, confirming its efficiency in real-time speech synthesis. These results highlight CSV-EHO’s effectiveness in generating high-quality, expressive, and computationally efficient speech.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Lightweight text-to-speech generation with convolutional swin transformer-based conditional variational Elk Herd Optimizer

  • Tarannum Shaikh,
  • Ashish Jadhav

摘要

Text-To-Speech (TTS) technology has progressive substantially, yet challenges persist in achieving human-like speech quality, emotional expressiveness, and efficient multilingual synthesis. To address these issues, this research proposes the Convolutional Swin transformer-based conditional Variational Elk Herd Optimizer (CSV-EHO) system, an advanced text-augmented speech synthesis model. The system integrates text normalization with GPT-3, phoneme conversion using an adapted Grapheme-to-Phoneme (G2P) model, and Mel-spectrogram generation leveraging Swin Transformer and convolutional networks. The Elk Herd Optimizer (EHO) enhances phoneme-to-spectrogram mapping, while a dual-path WaveNet vocoder with Cosine Wavelet Transform refines clarity and naturalness. Experimental evaluations on VCTK and LJ Speech datasets demonstrate excellent performance, achieving a MOS of 4.57 ± 0.21, CMOS of + 0.45, SSIM of 0.91, phoneme accuracy of 98.2%, and WER of 3.8%, along with PESQ scores of 2.9 and 2.6. The model also attains an inference time of 0.011 s, confirming its efficiency in real-time speech synthesis. These results highlight CSV-EHO’s effectiveness in generating high-quality, expressive, and computationally efficient speech.