The development of an accurate and contextually relevant translation system from Arabic text to Arabic gloss text presents significant challenges due to the complexity of Arabic grammar and the multidimensional nature of sign language. This paper investigates the impact of various data augmentation techniques on enhancing the performance of such translation systems. We leverage the ArSL dataset, a comprehensive bilingual parallel corpus in the health domain, and employ three distinct augmentation strategies: Blank Replacement, Synonyms Replacement, and Sentence Paraphrasing. Our findings demonstrate that data augmentation substantially improves the translation accuracy, as evidenced by the BLEU scores. The Data augmentation methods achieved the highest performance with a validation BLEU score of 92.77 and a test BLEU score of 92.15, significantly outperforming the non-augmented dataset. This study underscores the potential of data augmentation in overcoming dataset limitations and enhancing machine translation models for Arabic sign text.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Arabic Gloss Machine Translation Through Data Augmentation

  • Doaa Alghamdi,
  • Mansour Alsulaiman,
  • Yousef Alohali,
  • Mohamed A. Bencherif,
  • Mohammed Algabri,
  • Mohamed A. Mekhtiche

摘要

The development of an accurate and contextually relevant translation system from Arabic text to Arabic gloss text presents significant challenges due to the complexity of Arabic grammar and the multidimensional nature of sign language. This paper investigates the impact of various data augmentation techniques on enhancing the performance of such translation systems. We leverage the ArSL dataset, a comprehensive bilingual parallel corpus in the health domain, and employ three distinct augmentation strategies: Blank Replacement, Synonyms Replacement, and Sentence Paraphrasing. Our findings demonstrate that data augmentation substantially improves the translation accuracy, as evidenced by the BLEU scores. The Data augmentation methods achieved the highest performance with a validation BLEU score of 92.77 and a test BLEU score of 92.15, significantly outperforming the non-augmented dataset. This study underscores the potential of data augmentation in overcoming dataset limitations and enhancing machine translation models for Arabic sign text.