错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Natural-Sounding Speech to Text in English

  • Kriss Saulitis,
  • Evalds Urtans,
  • Vairis Caune

摘要

This study focuses on a systematic review of the literature and an experimental comparison of 20 English speech synthesis methods. Nine of the models were subjected to a quantitative analysis, using selected samples from the Common Voice data set and using criteria to assess both quality and precision. The research methodology includes the configuration of speech synthesis models to generate audio samples, which are then used to compare models based on established criteria. The NISQA model is used to evaluate speech quality through machine learning, mimicking the subjective MOS metric. Character and word error rate metrics are used to evaluate the precision of the synthesized samples. The CoMoSpeech model showed the best quality indicators (MOS - 3.85), while the VITS model demonstrated the highest precision (CER - 1.48%) and the total average of the metric (Synthesized samples and code repository are available at https://research.saulitis.dev/english-speech-synthesis-comparison-2024 ).