错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toxicity Detection and Classification in Arabic Text

  • Ahmed Abulohoom,
  • Ashraf Elnagar

摘要

Utilizing a corpus which allows for the study of diverse forms of the Arabic language, we have taken on the challenge of detecting and categorizing forms of toxic content in Arabic text. The corpus contains a large number of Arabic utterances which are tagged with binary labels for toxic/non-toxic content, and multi-class labels for toxic categories such as racism, violence, and hate speech. Our methodology contains a robust pipeline of data preprocessing which clears a path toward quality training and evaluation. We try different BERT-based models that have been selected due to the ability to process dialectal and modern standard Arabic. The MARBERTv2-based model revealed its superior ability in doing both binary and multi-classification. It achieved high F1-scores, indicating that it is effective in spite of all challenges presented by the corpus. From this study we conclude that natural language processing has proven its efficacy in identifying toxic Arabic language content. Furthermore, it paves the way to develop systems that detoxify Arabic text.