错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A New Document Representation Technique for Hate Speech Detection Using a New Term Weight Measure and Word Embedding Techniques

  • Raju Rdara,
  • N. Swapna,
  • Urmila Dara

摘要

Due to the advent of various Internet technologies in social media platforms like Twitter, Facebook, LinkedIn, Instagram, and review sites, people are changing their way of communication with other people. These platforms allow the people to freely share their knowledge, opinions to individual or group of people. Most of them used these platforms for welfare of others, but some of them misused these environments by spreading false information, hateful texts about a person, product, or any other entity. Hate speech (HS) is any expression, whether direct or indirect, that disparages another person or group of people based on their gender, sexual orientation, religion, ethnicity, or disability. The terminology of authors employ in their writings is crucial in determining whether or not a text contains hate speech. The weight of a term in a text also play prominent role to differentiate the hate speech from normal text. In this article, we proposed a new document representation method for hate speech detection. The proposed method represents the documents as vectors by using two varieties of information such as the document representation with the features identified through feature selection algorithm and the document vectors represented by using the word vectors generated through word embedding techniques. We developed a new term weight measure to compute the weight value of a vector and compared the performance of proposed term weight measure with different existing term weight measures. The text document vectors are trained with machine learning algorithm. Support vector machine (SVM) is used for producing the model. The SVM achieved 0.8964 score of accuracy for hate speech detection.