Talk like a Local: Evaluating Large Language Models for Arabic Dialect Translation Using Similarity Scores
摘要
This paper introduces a promising approach for evaluating the quality of translations between Modern Standard Arabic (MSA) and the Egyptian (Cairo) dialect using Large Language Models, specifically Claude and AraT5. We demonstrate how similarity scores provide a more meaningful and nuanced measure of translation quality, capturing semantic preservation and model performance improvements through fine-tuning on dialect-specific datasets. Our results highlight the potential of semantic similarity scores as a valuable evaluation tool for Arabic dialect translation using large language models. These findings suggest that semantic similarity scores could become a standard evaluation tool for Arabic dialect translations, enhancing the development and assessment of new high-quality datasets and advancing the Arabic NLP domain.