A Comparative Analysis of Models for Dark Web Data Classification
摘要
The Dark Web has raised significant concerns due to its covert nature and potential involvement in criminal activities. In this research study, we present a comprehensive examination of Dark Web data classification into distinct categories such as pornography, firearms, drugs, hacking, cryptocurrency, and others. To train our models, we employed a labeled CODA dataset and enhance it with unlabeled data collected through web crawling for testing purposes. To annotate the data, we utilize both Term Frequency-Inverse Document Frequency (TF-IDF) and Bag-of-Words (BOW) techniques [17–20]. We implement three classification models: Support Vector Machine (SVM), Naive Bayes (NB), and Recurrent Neural Network (RNN) to assess their effectiveness in categorizing the unlabeled test data. Through rigorous evaluation and comparison, we thoroughly analyze the performance of these models, offering valuable insights into the task of Dark Web data classification. This research makes a significant contribution to the development of improved methods for comprehending and monitoring illicit activities on the Dark Web, ultimately enhancing online safety and security.