<p>Considering the rapid growth of cyber threats and the increasing vulnerabilities of networked systems, ensuring their security has become one of the fundamental challenges. An intrusion detection system (IDS) plays a crucial role in protecting network data from misuse and suspicious activities by identifying malicious actions based on pre-defined patterns. However, traditional IDS models, which were primarily based on singular algorithms, have not been able to fully meet the expectations and still need further improvements. Many existing approaches have attempted to improve detection accuracy by employing a large number of models. While this has led to high accuracies, it has also increased system complexity and computational cost. Furthermore, high false-positive rates (FPRs) remain a significant issue, reducing the reliability of these models. Additionally, the presence of irrelevant and redundant features in datasets complicates the classification process. To address these challenges, this paper proposes a more efficient approach using the CICIDS2017 dataset. First, essential preprocessing steps, including data balancing and min–max normalization, are applied to improve data quality. Then, the mRMR (minimum redundancy maximum relevance) filter method is used to remove redundant and irrelevant features while selecting the most relevant ones. Finally, a stacking ensemble classification model is implemented, combining k-nearest neighbors (kNNs), decision tree (DT), and neural network (NN) as base learners, with random forest (RF) as the meta-learner, to enhance intrusion detection performance while maintaining a balance between accuracy and computational efficiency. Experimental results demonstrate that the proposed model outperformed singular algorithms by achieving an accuracy of 99.79%, while also reducing the false-positive rate (FPR) to 0.0002, which is significant. By effectively analyzing network traffic data, this approach enhances both detection accuracy and reliability, offering a promising direction for improving IDS.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multivariate filter feature selection and stacking-based ensemble learning algorithm for network intrusion detection

  • Seyedeh Maryam Hashemi Jouybari,
  • Alireza Janbaz,
  • Ahmad Esfandiari

摘要

Considering the rapid growth of cyber threats and the increasing vulnerabilities of networked systems, ensuring their security has become one of the fundamental challenges. An intrusion detection system (IDS) plays a crucial role in protecting network data from misuse and suspicious activities by identifying malicious actions based on pre-defined patterns. However, traditional IDS models, which were primarily based on singular algorithms, have not been able to fully meet the expectations and still need further improvements. Many existing approaches have attempted to improve detection accuracy by employing a large number of models. While this has led to high accuracies, it has also increased system complexity and computational cost. Furthermore, high false-positive rates (FPRs) remain a significant issue, reducing the reliability of these models. Additionally, the presence of irrelevant and redundant features in datasets complicates the classification process. To address these challenges, this paper proposes a more efficient approach using the CICIDS2017 dataset. First, essential preprocessing steps, including data balancing and min–max normalization, are applied to improve data quality. Then, the mRMR (minimum redundancy maximum relevance) filter method is used to remove redundant and irrelevant features while selecting the most relevant ones. Finally, a stacking ensemble classification model is implemented, combining k-nearest neighbors (kNNs), decision tree (DT), and neural network (NN) as base learners, with random forest (RF) as the meta-learner, to enhance intrusion detection performance while maintaining a balance between accuracy and computational efficiency. Experimental results demonstrate that the proposed model outperformed singular algorithms by achieving an accuracy of 99.79%, while also reducing the false-positive rate (FPR) to 0.0002, which is significant. By effectively analyzing network traffic data, this approach enhances both detection accuracy and reliability, offering a promising direction for improving IDS.