French-Fulfulde Textless and Cascading Speech Translation: Towards a Dual Architecture
摘要
Speech-to-Speech translation is attracting increasing attention from researchers due to its potential of easing the communication. It can be leveraged to develop voice user interfaces for important services, such as agriculture in low-literate communities where poorly resourced languages such as Fulfulde are used. If an earlier technique, such as cascading speech, requires textual data, a recent and textless one, such as direct speech-to-speech, only considers speech corpus, playing a crucial role in considering oral languages. In this work, a general approach for a dual architecture is proposed integrating direct speech-to-speech and speech-to-speech translation using automatic speech recognition, machine translation, and text-to-speech translation. Beyond proposing an architecture, this paper focuses on an important step in cascading speech translation, namely automatic speech recognition (speech to text translation). Automatic speech recognition for the Fulfulde language, using the Kaldi toolkit, allowed obtaining an average Word Error Rate of 28.91% for S2T with the monophone acoustic model and a WER of 26.58% for S2T with the triphone acoustic model. The dataset used is based on agricultural words recorded by natives of northern Cameroon.