Language resources and particularly the lack of computational lexicons is one of the central problems that face making progress in Moroccan Arabic NLP tasks. Another problem is resources reusability where it is difficult to employ the few existing resources in common NLP tasks. In this paper, we propose to alleviate these problems by building a reusable bi-lingual lexicon addressing both Moroccan Arabic and Arabic languages. For this purpose, we compiled data from different sources including printed and digital lexicons as well as speech and social media text. To meet users’ needs in a systematic and easily accessible way, the developed resource is structured following the Lexical Markup Framework standard and then hosted in software architecture with full respect to interoperability rules. Our lexicon contains almost 13000 lemmas with their Arabic equivalents and is manually annotated with useful metadata such as Part of Speech, Origin, and root. The lexicon is designed for practical NLP use such as morphological analysis and automatic translation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Compiling a Bilingual Lexicon Using a Semi-automatic Approach

  • Ridouane Tachicart,
  • Karim Bouzaoubaa,
  • Driss Namly

摘要

Language resources and particularly the lack of computational lexicons is one of the central problems that face making progress in Moroccan Arabic NLP tasks. Another problem is resources reusability where it is difficult to employ the few existing resources in common NLP tasks. In this paper, we propose to alleviate these problems by building a reusable bi-lingual lexicon addressing both Moroccan Arabic and Arabic languages. For this purpose, we compiled data from different sources including printed and digital lexicons as well as speech and social media text. To meet users’ needs in a systematic and easily accessible way, the developed resource is structured following the Lexical Markup Framework standard and then hosted in software architecture with full respect to interoperability rules. Our lexicon contains almost 13000 lemmas with their Arabic equivalents and is manually annotated with useful metadata such as Part of Speech, Origin, and root. The lexicon is designed for practical NLP use such as morphological analysis and automatic translation.