错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Classification of Offensive Tweet in Marathi Language Using Machine Learning Models

  • Archana Kumari,
  • Archana Garge,
  • Priyanshu Raj,
  • Gunjan Kumar,
  • Jyoti Prakash Singh,
  • Mohammad Alryalat

摘要

Offensive language identification is essential to make social media a safe and clean place to share one’s view. In this work, a model is proposed to automatically classify offensive tweets into offensive and not offensive classes of low-resource language. Marathi is spoken in Western India. Marathi being a low-resource language, lacks a comprehensive list of stopwords and proper stammer. To fill this gap, we created a list of stopwords for stopword removal and a list of suffixes to identify the root word in the Marathi language. Two different methods, Label Vectorizer and term frequency-inverse document frequency (TF-IDF) Vectorizer, are used to extract features from the text and then these features are used with six different conventional machine learning classifiers to classify a Marathi tweet into offensive or non-offensive.