A comprehensive review on detection of hate speech for multi-lingual data
摘要
The paper addresses the critical challenge of detecting and mitigating hate speech in Hindi across social media platforms, utilizing natural language processing (NLP) techniques. It underscores the complexities associated with the Hindi language, including the prevalence of code-switching, the existence of various dialects, and the inconsistent use of Romanized Hindi. These linguistic factors present significant challenges for developing automated systems capable of reliably identifying hate speech. The paper emphasizes the urgent need to confront hate speech in Hindi due to its potential to incite social unrest, propagate misinformation, and foster a culture of intolerance within India's diverse socio-political landscape. The research aims to address existing gaps in the detection of Hindi hate speech through employing advanced machine learning and deep learning methodologies. The overarching objective is to devise solutions that are not only technologically robust but also deeply attuned to the unique linguistic and cultural nuances of the Hindi language.