错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Abusive Speech Detection and Politeness Transfer

  • K. Preetham,
  • D. Arun Arumugham,
  • M. Yogesh Kumar,
  • B. Shameedha Begum

摘要

In the recent times of lockdown, the usage of all kinds of online platforms like Twitter, Facebook, YouTube and Reddit have increased by quite an extent. In addition to using these platforms for creating and sharing positive and inspiring content, a lot of hate and anger comments also seem to be prevalent in them. These problems are tackled by first detecting these forms of hate speech in textual data, then imparting “politeness” to the hateful comments while preserving the meaning conveyed. For the first phase of abusive speech detection, Baseline models like Logistic Regression, Naive Bayes, SVM, Random Forest and Decision Tree were trained and analyzed. Next, state-of-the-art models like LSTMs, Bi-LSTMs and Transformers were trained for classification of text. Word vectorization models like BOW and TF-IDF and also GloVe embeddings were used and evaluated on the models. It was found that Logistic Regression (with BOW), SVM (with TFIDF) and LSTMs were better performing than others in their categories. A hybrid model of the best performing classifiers was finally used. The next phase of politeness transfer to the abusive text was explored using BERT’s language model and its bidirectional property of understanding context to reduce the average toxicity of input sentences.