This paper evaluates the performance of different large language models (LLMs) in translating textual data from Saudi Arabic, a low-resource language, into English. In this investigation we employ the state-of-the-art language models namely; ChatGPT-4, Claude-3 and Palm-2. We assess the capabilities of these LLMs on the Arabic Semantic Textual Similarity (STS) dataset. The evaluation covers different aspects, including the standard evaluation metrics, prompt design, and comparison with baselines systems namely; Google Translator, QuillBot Translator and Systran Translator. We conducted human evaluation on the generated translation and analysis the most frequent translation error using our sample dataset and different models. Our findings reveal significant insights into the strengths of ChatGPT (GPT-4) model in handling and translating dialectal Arabic with the highest Bilingual Evaluation Understudy (BLEU) score among all participated models (46.56).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Performance of LLMs When Translating Saudi Arabic as Low Resource Language

  • Salwa Alahmari,
  • Eric Atwell,
  • Mohammad Alsalka,
  • Hadeel Saadany

摘要

This paper evaluates the performance of different large language models (LLMs) in translating textual data from Saudi Arabic, a low-resource language, into English. In this investigation we employ the state-of-the-art language models namely; ChatGPT-4, Claude-3 and Palm-2. We assess the capabilities of these LLMs on the Arabic Semantic Textual Similarity (STS) dataset. The evaluation covers different aspects, including the standard evaluation metrics, prompt design, and comparison with baselines systems namely; Google Translator, QuillBot Translator and Systran Translator. We conducted human evaluation on the generated translation and analysis the most frequent translation error using our sample dataset and different models. Our findings reveal significant insights into the strengths of ChatGPT (GPT-4) model in handling and translating dialectal Arabic with the highest Bilingual Evaluation Understudy (BLEU) score among all participated models (46.56).