Learning Arabic as a foreign language (L2) is becoming increasingly important in a rapidly expanding world of linguistic and cultural exchange. However, the development of effective educational tools and appropriate teaching resources remains limited by the availability of quality data. The creation and expansion of dedicated textual corpora are therefore major challenges for improving the accessibility and relevance of educational resources. In this context, our work aims to enrich the educational data gathered from the “Global Language Online Support System platform” (GLOSS), specifically designed for foreign language learning. The aim is to develop a data augmentation methodology that generates synthetic texts close to the originals, both linguistically and pedagogically and this using LLM approaches. To achieve this, we have designed a complete pipeline incorporating pre-processing, morphological annotation and evaluation stages. The use of the AraBERT model, combined with similarity metrics, guarantees the quality and consistency of the texts generated. This approach aims not only to increase the resources available, but also to support a variety of applications such as readability assessment and the development of advanced pedagogical tools.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enriching Arabic Educational Data with AraBERT and Similarity Assessment

  • Naoual Nassiri,
  • Ayoub Jibouni,
  • Sahar Saoud

摘要

Learning Arabic as a foreign language (L2) is becoming increasingly important in a rapidly expanding world of linguistic and cultural exchange. However, the development of effective educational tools and appropriate teaching resources remains limited by the availability of quality data. The creation and expansion of dedicated textual corpora are therefore major challenges for improving the accessibility and relevance of educational resources. In this context, our work aims to enrich the educational data gathered from the “Global Language Online Support System platform” (GLOSS), specifically designed for foreign language learning. The aim is to develop a data augmentation methodology that generates synthetic texts close to the originals, both linguistically and pedagogically and this using LLM approaches. To achieve this, we have designed a complete pipeline incorporating pre-processing, morphological annotation and evaluation stages. The use of the AraBERT model, combined with similarity metrics, guarantees the quality and consistency of the texts generated. This approach aims not only to increase the resources available, but also to support a variety of applications such as readability assessment and the development of advanced pedagogical tools.