错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CFMD: Corpus for Moroccan Dialect as Under Researched Dialect

  • Hajar Zaidani,
  • Abderrahim Maizate,
  • Mohammed Ouzzif,
  • Rim Koulali

摘要

The rise of social media has revolutionised numerous fields within artificial intelligence, and one of the domains greatly impacted is natural language processing. With the widespread use of text, audio, and video in social media platforms, there is a unique opportunity to delve into the intricacies of lesser-documented dialects, including the Moroccan dialect. However, due to the scarcity of readily available datasets for this specific dialect, we recognized the need to create our own corpus to facilitate research and development in this area. The Moroccan dialect, being primarily a spoken language with a lack of standardized rules and a rich variety of subdialects, presents a challenge in terms of data collection and analysis. In order to overcome this obstacle, we devised an innovative approach leveraging the vast content available on the YouTube platform. By utilising the YouTube platform, we were able to download a wide range of audio content. To ensure accuracy and reliability we enlisted native speakers for accurate transcription. The resulting corpus offers a rich collection of texts that explore a diverse range of topics. Furthermore; this corpus serves as a valuable resource for the development and improvement of machine learning algorithms and models for the Moroccan dialect.