<p>Hostile social media content has become a considerable problem for people, governments, and organizations due to the internet’s expanding reach. Using an automated approach, such material must be recognized from massive amounts of data. However, because of the requirement for more appropriate datasets, most studies have concentrated on English texts, leaving a noticeable void in other low-resource languages. Hindi’s varied syntactic structure makes it challenging to detect hostility. Hence, it is essential to create efficient detection techniques for low-resource languages. In this study, we have used a dataset written in the Hindi Devanagari script, which consists of approximately 8300 instances and covers five distinct groups of hostile and non-hostile incidents. We employ attention-based pre-trained models fine-tuned on Hindi data to leverage this novel resource. We propose a hybrid model combining Convolutional Neural Network and Decision Tree for effective content classification and analysis. This integrated approach uses the feature extraction capabilities of CNN with the interpretability and decision-making structure of the Decision Tree for accurately classifying complex content. Using this novel hybrid approach, we have achieved notable performance by achieving an overall accuracy of 98.98% and an F1 score of 98.93%. It demonstrates applying advanced natural language processing techniques for multi-class text classification in Hindi.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel ensemble of convolutional neural network with decision tree classifier for hostile post detection in hindi

  • Santosh Rajak,
  • Ujwala Baruah

摘要

Hostile social media content has become a considerable problem for people, governments, and organizations due to the internet’s expanding reach. Using an automated approach, such material must be recognized from massive amounts of data. However, because of the requirement for more appropriate datasets, most studies have concentrated on English texts, leaving a noticeable void in other low-resource languages. Hindi’s varied syntactic structure makes it challenging to detect hostility. Hence, it is essential to create efficient detection techniques for low-resource languages. In this study, we have used a dataset written in the Hindi Devanagari script, which consists of approximately 8300 instances and covers five distinct groups of hostile and non-hostile incidents. We employ attention-based pre-trained models fine-tuned on Hindi data to leverage this novel resource. We propose a hybrid model combining Convolutional Neural Network and Decision Tree for effective content classification and analysis. This integrated approach uses the feature extraction capabilities of CNN with the interpretability and decision-making structure of the Decision Tree for accurately classifying complex content. Using this novel hybrid approach, we have achieved notable performance by achieving an overall accuracy of 98.98% and an F1 score of 98.93%. It demonstrates applying advanced natural language processing techniques for multi-class text classification in Hindi.