Disfluent-to-Fluent Tunisian Dialect Speech Translation with Fine-Tuning Pre-trained Language Models
摘要
In contrast to written texts and prepared speeches, conversational/spontaneous speech has a very high degree of freedom and includes a huge number of disfluencies. Detecting disfluencies using transformer-based models has advanced state-of-the-art performance. In this work, we aim to process disfluencies in the spontaneous tunisian dialect speech by generating fluent utterances from disfluent transcripts. We propose a transformer-based model by fine-tuning the pre-trained T5 language model. Using this model, we achieved an F-Measure score of 74,71% based on the evaluation data set part of DisCoTAT.