Detecting and addressing covert hate speech poses significant challenges in online platforms where discriminatory messages are often disguised within seemingly innocuous content. In this study, we investigate the effectiveness of contextual analysis in bypassing the nuances of covert hate speech. Our research explores the impact of prompt engineering and context addition on the classification of overt and covert hate speech across diverse target groups, including Roma, migrants, LGBTQ+, and individuals of African descent. Through experimental trials using generative models like GPT-3.5 and GPT-4, our findings reveal that the addition of context, not only improves the overall performance of the models in general hate speech (from 75.0% to 79.6% F1 score, for the positive class), but also significantly improves the classification of covert hate speech, increasing True Positives by 21.64% (absolute) compared to the 6.5% in overt hate speech. Despite these improvements, the addition of context also increased the number of False Positives, indicating that a further refinement is needed for this contextual analysis. Moreover, target group analysis demonstrates a correlation between the prevalence of covert hate speech and model performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bypassing the Nuances of Portuguese Covert Hate Speech Through Contextual Analysis

  • Gil Ramos,
  • Fernando Batista,
  • Ricardo Ribeiro,
  • Pedro Fialho,
  • Sérgio Moro,
  • António Fonseca,
  • Rita Guerra,
  • Paula Carvalho,
  • Catarina Marques,
  • Cláudia Silva

摘要

Detecting and addressing covert hate speech poses significant challenges in online platforms where discriminatory messages are often disguised within seemingly innocuous content. In this study, we investigate the effectiveness of contextual analysis in bypassing the nuances of covert hate speech. Our research explores the impact of prompt engineering and context addition on the classification of overt and covert hate speech across diverse target groups, including Roma, migrants, LGBTQ+, and individuals of African descent. Through experimental trials using generative models like GPT-3.5 and GPT-4, our findings reveal that the addition of context, not only improves the overall performance of the models in general hate speech (from 75.0% to 79.6% F1 score, for the positive class), but also significantly improves the classification of covert hate speech, increasing True Positives by 21.64% (absolute) compared to the 6.5% in overt hate speech. Despite these improvements, the addition of context also increased the number of False Positives, indicating that a further refinement is needed for this contextual analysis. Moreover, target group analysis demonstrates a correlation between the prevalence of covert hate speech and model performance.