<p>Illegal marketplaces have increasingly migrated to the deep/dark web and high-risk social platforms, facilitating the anonymous trade of drugs, weapons, and stolen credentials. Detecting and categorizing such content remains challenging due to scarce labeled data, rapidly evolving language, diverse structural characteristics, and the limitations of existing approaches in handling noisy, cross-platform illicit content. To address these challenges, we propose a novel two-stage hierarchical semi-supervised framework that integrates domain-adapted ModernBERT embeddings, manually engineered structural features, and an entropy-based ensemble learning strategy for illicit marketplace detection and classification. ModernBERT is fine-tuned on domain-specific data to better capture specialized jargon, obfuscated expressions, and long-context linguistic patterns commonly found in illicit marketplaces. These embeddings are combined with layout, pattern-specific, and metadata features to enrich document representation across heterogeneous platforms. In the first stage, sales-related documents are identified using an ensemble of XGBoost, Random Forest, and SVM classifiers within a self-training framework enhanced by a novel entropy-based weighted voting mechanism that dynamically adjusts classifier contributions based on prediction confidence. In the second stage, three specialized semi-supervised XGBoost classifiers categorize detected sales content into drug, weapon, and stolen credential sales. To develop and evaluate the proposed framework, a 21,575-sample multi-source corpus comprising 1,575 labeled and 20,000 unlabeled samples is collected from the deep/dark web, Telegram, Reddit, and Pastebin. Experimental results demonstrate superior performance, achieving macro-averaged scores of 0.96489 accuracy, 0.93467 F1, and 0.95388 TMCC. Also, the proposed framework consistently improves over baseline models, including ModernBERT, BERT, Longformer, ALBERT, BigBird, and DarkBERT, with accuracy gains of 3.1-6.7%, F1 gains of 6.4-15.6%, and TMCC gains of 4.1-10.0%. Furthermore, experiments on the DUTA and CoDA datasets confirm the robustness, generalizability, and practical effectiveness of the proposed framework for real-world illicit marketplace detection and classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A language model-driven semi-supervised ensemble framework for illicit market detection across deep/dark web and social platforms

  • Navid Yazdanjue,
  • Morteza Rakhshaninejad,
  • Hossein Yazdanjouei,
  • Mohammad Sadegh Khorshidi,
  • Mikko S. Niemelä,
  • Fang Chen,
  • Amir H. Gandomi

摘要

Illegal marketplaces have increasingly migrated to the deep/dark web and high-risk social platforms, facilitating the anonymous trade of drugs, weapons, and stolen credentials. Detecting and categorizing such content remains challenging due to scarce labeled data, rapidly evolving language, diverse structural characteristics, and the limitations of existing approaches in handling noisy, cross-platform illicit content. To address these challenges, we propose a novel two-stage hierarchical semi-supervised framework that integrates domain-adapted ModernBERT embeddings, manually engineered structural features, and an entropy-based ensemble learning strategy for illicit marketplace detection and classification. ModernBERT is fine-tuned on domain-specific data to better capture specialized jargon, obfuscated expressions, and long-context linguistic patterns commonly found in illicit marketplaces. These embeddings are combined with layout, pattern-specific, and metadata features to enrich document representation across heterogeneous platforms. In the first stage, sales-related documents are identified using an ensemble of XGBoost, Random Forest, and SVM classifiers within a self-training framework enhanced by a novel entropy-based weighted voting mechanism that dynamically adjusts classifier contributions based on prediction confidence. In the second stage, three specialized semi-supervised XGBoost classifiers categorize detected sales content into drug, weapon, and stolen credential sales. To develop and evaluate the proposed framework, a 21,575-sample multi-source corpus comprising 1,575 labeled and 20,000 unlabeled samples is collected from the deep/dark web, Telegram, Reddit, and Pastebin. Experimental results demonstrate superior performance, achieving macro-averaged scores of 0.96489 accuracy, 0.93467 F1, and 0.95388 TMCC. Also, the proposed framework consistently improves over baseline models, including ModernBERT, BERT, Longformer, ALBERT, BigBird, and DarkBERT, with accuracy gains of 3.1-6.7%, F1 gains of 6.4-15.6%, and TMCC gains of 4.1-10.0%. Furthermore, experiments on the DUTA and CoDA datasets confirm the robustness, generalizability, and practical effectiveness of the proposed framework for real-world illicit marketplace detection and classification.