<p>The present study develops a novel empirical framework to evaluate and compare the impact of eight different filter-based feature selection (FS) techniques and four wrapper-based FS techniques on the performance of four black-box models: Random Forest, Multilayer Perceptron, eXtreme Gradient Boosting, and Support Vector Machines using four intrusion datasets: CIC-IDS2018, CIC-ToN-IoT, NF-UNSW-NB15-v2, and NF-UQ-NIDS-v2. Moreover, the application of ensemble FS was introduced, leveraging the strengths of single FS techniques through three distinct strategies (intersection-based, multi-intersection-based, and fusion-based single techniques). These strategies were evaluated and compared using the Scott–Knott analysis and Borda count ranking technique to determine their effectiveness in enhancing FS processes. In this experiment, 288 variants of classifiers are evaluated to determine the features that positively impact the classification efficiency of all the cyber-attack scenarios used. Furthermore, due to the opacity of the classifiers used, the Global Surrogate (GS) interpretability technique has been used to determine the accuracy–interpretability trade-off. Experiments have demonstrated that Recursive Feature Elimination and Boruta FS techniques perform effectively across most datasets, achieving strong results when combined with ensemble models like RF and XGB, even when trained with a small number of features. However, consistency-based FS trained with ensemble models was the best choice when considering the number of features. Moreover, filter-based ensembles generally outperformed their singles, whereas in the case of wrappers, singles had significantly better results than their ensembles. Furthermore, regardless of the number of features, the fusion and multi-intersection of filters achieved the best overall performance over all datasets. However, taking into account both performance and the number of features, wrapper ensembles showed superior performance over all datasets. Finally, the findings showed the potential of the GS method to deal with the accuracy–interpretability trade-off for all opaque models and highlight its role in enhancing the transparency of all black-box models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature selection and global interpretability of black-box classification for intrusion detection

  • Houssam Zouhri,
  • Ali Idri

摘要

The present study develops a novel empirical framework to evaluate and compare the impact of eight different filter-based feature selection (FS) techniques and four wrapper-based FS techniques on the performance of four black-box models: Random Forest, Multilayer Perceptron, eXtreme Gradient Boosting, and Support Vector Machines using four intrusion datasets: CIC-IDS2018, CIC-ToN-IoT, NF-UNSW-NB15-v2, and NF-UQ-NIDS-v2. Moreover, the application of ensemble FS was introduced, leveraging the strengths of single FS techniques through three distinct strategies (intersection-based, multi-intersection-based, and fusion-based single techniques). These strategies were evaluated and compared using the Scott–Knott analysis and Borda count ranking technique to determine their effectiveness in enhancing FS processes. In this experiment, 288 variants of classifiers are evaluated to determine the features that positively impact the classification efficiency of all the cyber-attack scenarios used. Furthermore, due to the opacity of the classifiers used, the Global Surrogate (GS) interpretability technique has been used to determine the accuracy–interpretability trade-off. Experiments have demonstrated that Recursive Feature Elimination and Boruta FS techniques perform effectively across most datasets, achieving strong results when combined with ensemble models like RF and XGB, even when trained with a small number of features. However, consistency-based FS trained with ensemble models was the best choice when considering the number of features. Moreover, filter-based ensembles generally outperformed their singles, whereas in the case of wrappers, singles had significantly better results than their ensembles. Furthermore, regardless of the number of features, the fusion and multi-intersection of filters achieved the best overall performance over all datasets. However, taking into account both performance and the number of features, wrapper ensembles showed superior performance over all datasets. Finally, the findings showed the potential of the GS method to deal with the accuracy–interpretability trade-off for all opaque models and highlight its role in enhancing the transparency of all black-box models.