In today’s digital era, the surge of hate speech on online platforms has emerged as a significant societal issue. This paper addresses the requirement for robust hate speech detection systems, particularly in a language like Hindi, where research in this domain remains scarce despite the language's widespread use. We delve into the complexities of defining hate speech, its interpretation on various social media platforms, and the wide range of categories it includes. We present a systematic pipeline for hate speech detection, covering data collection and preprocessing, feature engineering, model training, and evaluation. Moreover, we discuss commonly used state-of-the-art models for identifying hate speech, such as BERT and its variations, CNN, LSTM, and others, along with feature engineering techniques that include N-grams, BoW, and TF-IDF. This study highlights the unique challenges of detecting hate speech in Hindi, a language with diverse dialects and informal expressions, often underexplored in existing research. We benchmark three models—BERT, CNN, and LSTM—on a diverse Hindi dataset and demonstrate the efficacy of state-of-the-art feature engineering techniques like N-grams and TF-IDF. Our findings pave the way for integrating multimodal approaches and expanding the scope of Hindi-specific hate speech detection, contributing to safer digital environments for millions of Hindi-speaking users.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Deep Learning Models for Hate Speech Detection in Hindi Textual Dataset

  • Rachna Narula,
  • Poonam Chaudhary

摘要

In today’s digital era, the surge of hate speech on online platforms has emerged as a significant societal issue. This paper addresses the requirement for robust hate speech detection systems, particularly in a language like Hindi, where research in this domain remains scarce despite the language's widespread use. We delve into the complexities of defining hate speech, its interpretation on various social media platforms, and the wide range of categories it includes. We present a systematic pipeline for hate speech detection, covering data collection and preprocessing, feature engineering, model training, and evaluation. Moreover, we discuss commonly used state-of-the-art models for identifying hate speech, such as BERT and its variations, CNN, LSTM, and others, along with feature engineering techniques that include N-grams, BoW, and TF-IDF. This study highlights the unique challenges of detecting hate speech in Hindi, a language with diverse dialects and informal expressions, often underexplored in existing research. We benchmark three models—BERT, CNN, and LSTM—on a diverse Hindi dataset and demonstrate the efficacy of state-of-the-art feature engineering techniques like N-grams and TF-IDF. Our findings pave the way for integrating multimodal approaches and expanding the scope of Hindi-specific hate speech detection, contributing to safer digital environments for millions of Hindi-speaking users.