<p>Online social networks have become the most popular medium for interpersonal communication, emotional expression, and information exchange. Leveraging artificial intelligence technologies, this study aims to analyze and address the issues of toxicity and abusiveness prevalent on X (formerly known as Twitter), specifically within Arabic-speaking communities. Data for this research was collected using the X API and meticulously annotated by native Arabic speakers to create a balanced dataset of toxic and non-toxic tweets. These tweets were categorized across various topics, including politics, sports, economics, religion, technology, and more. We employed topic analysis, dialect identification, sentiment analysis, and data frequency distribution techniques to examine the dataset. Our findings indicate that topics related to politics, religion, and pornography contain the highest levels of toxicity. Additionally, dialect identification reveals significant regional variations, with the most toxic content appearing in the Libyan, Egyptian, and Yemeni dialects. Sentiment analysis shows a predominance of negative sentiments in toxic tweets. Moreover, the data frequency distribution highlights commonly used abusive terms and phrases. This study underscores the critical need to address toxicity on social media platforms across different languages and cultures, providing valuable insights into the nature and distribution of harmful content in Arabic-speaking regions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing toxicity in Arabic social media: a study of regional dialects, sentiments, and toxic topics on X/Twitter

  • Loay Hatem,
  • Ahmed Omar,
  • Heba Mamdouh Farghaly,
  • Abdelmgeid A. Ali

摘要

Online social networks have become the most popular medium for interpersonal communication, emotional expression, and information exchange. Leveraging artificial intelligence technologies, this study aims to analyze and address the issues of toxicity and abusiveness prevalent on X (formerly known as Twitter), specifically within Arabic-speaking communities. Data for this research was collected using the X API and meticulously annotated by native Arabic speakers to create a balanced dataset of toxic and non-toxic tweets. These tweets were categorized across various topics, including politics, sports, economics, religion, technology, and more. We employed topic analysis, dialect identification, sentiment analysis, and data frequency distribution techniques to examine the dataset. Our findings indicate that topics related to politics, religion, and pornography contain the highest levels of toxicity. Additionally, dialect identification reveals significant regional variations, with the most toxic content appearing in the Libyan, Egyptian, and Yemeni dialects. Sentiment analysis shows a predominance of negative sentiments in toxic tweets. Moreover, the data frequency distribution highlights commonly used abusive terms and phrases. This study underscores the critical need to address toxicity on social media platforms across different languages and cultures, providing valuable insights into the nature and distribution of harmful content in Arabic-speaking regions.