Advancing offensive language detection in Arabic social media: a BERT-based ensemble learning approach
摘要
The growing ubiquity of online anonymity has significantly transformed the dynamics of participation and collaboration across digital platforms, especially through social media. While this veil of anonymity enables the free expression of personal opinions, it also leads to the spread of negative content, including offensive language. The detection and mitigation of such offensive language pose a considerable challenge and represent a dynamic research area. This paper introduces and evaluates a new ensemble learning approach designed for the Arabic offensive language detection task. The proposed approach combines three distinct models: the pretrained Bidirectional Encoder Representations from Transformers (BERT) model, BERT combined with Global Average and Global Max layers, and BERT augmented with pooled stacked Bidirectional Long Short-Term Memory (Bi-LSTMs). Outperforming the performance of the baseline OffensEval2020 winner in “SemEval-2020 Task 12 on Multilingual Offensive Language Identification in Social Media”, our model achieved a 90.97% F1-score on the original Arabic OffensEval2020 dataset and an enhanced 94.56% F1-score on the augmented dataset.