错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BERT and LLM-Based Multivariate Hate Speech Detection on Twitter: Comparative Analysis and Superior Performance

  • Xiaohou Shi,
  • Jiahao Liu,
  • Yaqi Song

摘要

The detection of toxic and hate speech in online social media is becoming increasingly necessary due to its prevalence and the potentially harmful consequences it can cause. Previous research has demonstrated the vital role that machine learning and natural language processing models have in identifying inappropriate language. In this study, the aim is to assess the viability of BERT for accurately predicting multivariate classifications related to hate speech on Twitter. The analysis will be conducted using the Twitter hate speech dataset. BERT has demonstrated exceptional performance in numerous areas of NLP, making it a potentially superior alternative to traditional machine learning approaches. Experiments were performed on the same dataset using 1-layer BERT, 2-layers BERT, and logistic regression models for both training and prediction purposes. The results demonstrate that the 2-layer BERT produces an accuracy of 85%. Additionally, we incorporated transfer learning techniques by leveraging a Large Language model GPT-3 and data augmentation strategies to further enhance model performance. This experiment reached a higher accuracy of 88%. As this is a multivariate classification problem with an asymmetrical dataset, we anticipate BERT and GPT-3 will achieve greater accuracy for the binary classification problem of identifying hate speech. These findings enhance the comprehension of hate speech detection in online material and the implications of various modeling approaches.