Detection of Cyberbullying Text in Bangla Using N-Gram Analysis and Machine Learning Approaches
摘要
Facebook, Twitter, Instagram, and YouTube are the most widely accepted communities or social media in these present times. Users, from tiny toddlers to teenagers to adults, are responsible for increasing the attraction of these social media. Cyberbullying can produce the effect, anguish, disappointment, and harassment that has piqued the interest of many individuals. Cyberbullying detection on Bangla text is a new research area that is gaining popularity, although it faces certain problems due to the restricted resources available. While a variety of systems have been used to classify English texts, there have been comparatively few studies on Bangla text classification. In this research, we offer an automatic cyberbullying detection model based on gram analysis model and machine learning. This research aims to detect and classify bullying words using several classification algorithms like Logistic Regression, Decision Tree Classifier, Random Forest Classifier, Multinomial NB, K Neighbors Classifier, Linear SVM, Radial Basis Function SVM, SGD, XGB Classifier. Here, a total of 1088 records or comments are scrutinized. There is after, pre-processing techniques, word tagging like general tagging, and abusive word tagging are used to extract features and construct the structure of comments. This research is then carried out utilizing several machine learning classifier algorithms and the Gram model. Thus, this paper shows the highest performance of accuracy with Unigram, Bigram, and Trigram which is achieved at 90.95, 88.99, and 86.96%. In which Multinomial NB as a classifier gives maximum accuracy value for Unigram and Bigram.