Overcoming Linguistic Challenges: Leveraging Deep Learning Techniques for Abusive Text Detection in Assamese Language
摘要
Social media has become a fundamental tool for communication in global community and idea exchange but can negatively affect the experience of users and raise conflicts online. With this unprecedented rise in internet usage, abusive content on social media has increased, hence requiring the development of techniques to detect such content effectively. This paper aims to identify abusive text in Assamese, one of the Indo-Aryan languages spoken in northeastern India. The language is marked as complex due to many dialectical modifications, incorporation of Hindi and English through code-mixing, and writing styles that are often irregular. This paper tries to compare Convolutional Neural Networks (CNN) and Gated Recurrent Units (GRU) in determining abusive comments on social media text in Assamese. The author has used a carefully selected dataset comprising 2669 sentences spread over multiple social media platforms. The GRU model demonstrated better performance than the CNN model with 89.23% accuracy. This work is thereby contributing to natural language processing for low-resource scenarios and deep learning algorithms to identify abusive language for low-resource languages.