<p>This systematic review, conducted under the PRISMA 2020 framework, investigates the application of Natural Language Processing (NLP) in insider threat detection, integrating Machine Learning (ML) and Deep Learning (DL) techniques, based on literature from 2019 to 2024. Addressing research questions (RQ1–RQ5) on techniques, datasets, performance metrics, challenges, and future directions, and evaluated via quality assessment criteria (QA1–QA4), the study synthesizes 66 high-quality studies from an initial 132 records across databases like IEEE, ScienceDirect, and Scopus. Findings reveal a surge in publications, peaking in 2023 and 2024, with Deep Learning (30.3%), hybrid NLP-ML/DL (27.3%), traditional ML (27.3%), and NLP-only (15.2%) approaches dominating, leveraging datasets like CERT and Enron to achieve accuracies up to 99.2% and F1-scores above 94%. However, reliance on synthetic data limits real-world applicability. Challenges include model complexity, explainability issues, binary classification biases, and dataset gaps, while future opportunities lie in lightweight, interpretable models, hybrid pipelines, anonymized datasets, and digital twins. The review highlights NLP’s transformative potential in enhancing cybersecurity against insider risks, offering a novel taxonomy and critique. Interdisciplinary efforts are recommended to align academic innovations with practical deployment, addressing gaps in real-time analysis and dataset diversity to strengthen organizational security frameworks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A systematic review on insider threat detection using natural language processing

  • Ketan Kundiya,
  • Yashodhara Haribhakta

摘要

This systematic review, conducted under the PRISMA 2020 framework, investigates the application of Natural Language Processing (NLP) in insider threat detection, integrating Machine Learning (ML) and Deep Learning (DL) techniques, based on literature from 2019 to 2024. Addressing research questions (RQ1–RQ5) on techniques, datasets, performance metrics, challenges, and future directions, and evaluated via quality assessment criteria (QA1–QA4), the study synthesizes 66 high-quality studies from an initial 132 records across databases like IEEE, ScienceDirect, and Scopus. Findings reveal a surge in publications, peaking in 2023 and 2024, with Deep Learning (30.3%), hybrid NLP-ML/DL (27.3%), traditional ML (27.3%), and NLP-only (15.2%) approaches dominating, leveraging datasets like CERT and Enron to achieve accuracies up to 99.2% and F1-scores above 94%. However, reliance on synthetic data limits real-world applicability. Challenges include model complexity, explainability issues, binary classification biases, and dataset gaps, while future opportunities lie in lightweight, interpretable models, hybrid pipelines, anonymized datasets, and digital twins. The review highlights NLP’s transformative potential in enhancing cybersecurity against insider risks, offering a novel taxonomy and critique. Interdisciplinary efforts are recommended to align academic innovations with practical deployment, addressing gaps in real-time analysis and dataset diversity to strengthen organizational security frameworks.