Twitter Trolling Detection Using Machine Learning
摘要
The increased usage of media has resulted in an increase, in cyberbullying and online trolling, and this study aims to identify and address trolling behavior through the application of machine learning algorithms also the research incorporates seven algorithms, Decision Tree, KNN, Logistic Regression, Naive Bayes Random Forest, Simple ANN and SVM. Additionally, four feature extraction techniques are employed: Bag-of-Words, TF IDF, Word2Vec and Glove. The analysis returned 28 combinations, with Glove word embeddings attaining the best accuracy at 89.77%. Bag-of-Words fared ideally when partnered with Logistic Regression and SVM, whereas TF-IDF excelled with Simple ANN and Naive Bayes. Word2Vec exhibited proficiency alongside Decision Tree and KNN. The Twitter dataset, characterized by unorthodox linguistic formulations, changes in syntax, and liberal use of slang, eludes typical text preprocessing techniques. The dataset also confronts us with symbols replacing letters, making typical symbol removal ineffective. This research underlines the relevance of customizing techniques to unique data sources and setting the way for future efforts to increase trolling detection.