错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tigrinya End-to-End Speech Recognition: A Hybrid Connectionist Temporal Classification-Attention Approach

  • Bereket Desbele Ghebregiorgis,
  • Yonatan Yosef Tekle,
  • Mebrahtu Fisshaye Kidane,
  • Mussie Kaleab Keleta,
  • Rutta Fissehatsion Ghebraeb,
  • Daniel Tesfai Gebretatios

摘要

The latest improvements in end-to-end Automatic Speech Recognition (ASR) systems have achieved outstanding results and have thus enabled the creation of state-of-the-art models for well-resourced languages. However, most languages, such as Tigrinya, are under-resourced, discouraging field efforts. Tigrinya is a Semitic language with over nine million speakers. This paper presents the first hybrid Connectionist Temporal Classification (CTC) with an attention-based end-to-end speaker-independent ASR model for Tigrinya. This initiative constructed new text and speech corpora encompassing multiple domains and thorough pre-processing, which amounted to about 170,000 phrases and sentences of text and 30 h of speech corpus. Data augmentation was applied to generate synthetic data for better generalization capability. A Recurrent Neural Network Language Model (RNN-LM) was also used for post-processing to complement the model to achieve even better results. Multiple experiments were conducted with different settings and parameters. Whilst keeping the data size/split constant and employing various combinations of data augmentation techniques along with varying LM’s vocabulary size showed improved performances, increasing the vocabulary size from 5k to 20k resulted in minute decoding improvement. Our best model exhibited a Character Error Rate (CER) of 14.28% and a Word Error Rate (WER) of 36.01%, which is significant considering this end-to-end approach is the first of its kind for the under-resourced Tigrinya language.