错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated Offensive Comment Detection for the Romanian Language

  • Andrei Paraschiv,
  • Andreea Cojocaru,
  • Mihai Dascalu

摘要

Offensive language can lead to uncomfortable situations, psychological harm, and, in particular cases, even violence. Social networks and websites struggle to reduce the prevalence of these messages by using an automated detector. One goal of Human-computer interaction (HCI) sciences is to provide respectful, safe, and user-friendly systems. This extends to any form of computer-mediated social interaction. This chapter contributes to this objective by proposing a Romanian language dataset for offensive message detection. We manually annotated 4,052 comments on a Romanian local news website into one of the following classes: non-offensive, targeted insults, racist, homophobic, and sexist. In addition, we establish a baseline of five automated classifiers, out of which the model based on RoBERT and two layers of CNN achieves the highest performance with a weighted F1-score of 74.74%.