<p>As the integration of Artificial Intelligence into network intrusion detection systems matures, a critical gap remains in the rigorous empirical benchmarking of datasets, preprocessing techniques, and model effectiveness. This article presents a comparative experimental study of anomaly and threat detection techniques used in network analysis through a multistep pipeline. First, we perform a structured comparison and Exploratory Data Analysis of the most commonly used network security datasets to quantitatively assess their balance, diversity, and real-world representativeness. Subsequently, we experimentally evaluate the performance of distinct Machine Learning, Deep Learning, and Hybrid Models derived under standardized preprocessing techniques to determine the most effective combinations for specific attack vectors. Our results demonstrate that hybrid architectures achieve superior generalisation, yet face challenges regarding computational overhead and cross-dataset adaptability. Addressing these limitations, we propose a proof-of-concept adaptive architecture designed to handle concept drift and adversarial threats. Finally, we outline a roadmap for future research, emphasizing the necessity of dynamic, verifiable AI-driven systems that operate reliably in evolving cybersecurity environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards maintainable AI-driven network anomaly and threat detection: a comparative analysis of datasets, preprocessing techniques, and model trade-offs

  • Antonio Lara-Gutierrez,
  • Carmen Fernandez-Gago,
  • Jose A. Onieva

摘要

As the integration of Artificial Intelligence into network intrusion detection systems matures, a critical gap remains in the rigorous empirical benchmarking of datasets, preprocessing techniques, and model effectiveness. This article presents a comparative experimental study of anomaly and threat detection techniques used in network analysis through a multistep pipeline. First, we perform a structured comparison and Exploratory Data Analysis of the most commonly used network security datasets to quantitatively assess their balance, diversity, and real-world representativeness. Subsequently, we experimentally evaluate the performance of distinct Machine Learning, Deep Learning, and Hybrid Models derived under standardized preprocessing techniques to determine the most effective combinations for specific attack vectors. Our results demonstrate that hybrid architectures achieve superior generalisation, yet face challenges regarding computational overhead and cross-dataset adaptability. Addressing these limitations, we propose a proof-of-concept adaptive architecture designed to handle concept drift and adversarial threats. Finally, we outline a roadmap for future research, emphasizing the necessity of dynamic, verifiable AI-driven systems that operate reliably in evolving cybersecurity environments.