<p>The adoption of Internet of Things (IoT) technologies has introduced significant security challenges due to the diversity and number of connected devices. Although machine learning (ML) is widely used in intrusion detection systems (IDSs), it does not always ensure a good trade-off between bias and variance. This study presents a large-scale evaluation of ensemble learning methods, focusing on boosting techniques while considering feature types and transformation strategies. Five boosting algorithms (adaptive boosting (ADB), gradient boosting machines (GBM), category boosting (CATB), light gradient boosting machine (LGBM), and extreme gradient boosting (XGB)) were evaluated using two transformation strategies and five filter-based feature selection (FS) methods: analysis of variance (ANOVA), Kendall’s tau, mutual information (MI), maximum relevance minimum redundancy (mRMR), and Chi-square (Chi2). Experiments were conducted on two large NetFlow-based IoT datasets (NF-ToN-IoT-v2 and NF-BoT-IoT-v2). Each model was assessed using Matthews correlation coefficient (MCC), Cohen’s kappa, F1-score, and accuracy, and ranked using the Scott–Knott statistical test and Borda count voting system. A total of 1160 models were trained using the Toubkal Supercomputer, whose high-performance computing (HPC) infrastructure enabled scalable parallel and distributed processing. The best results were obtained using XGB with 200 estimators and 18 selected features under the standardization transformation. For the NF-ToN-IoT-v2 dataset, the combination of ANOVA and mRMR yielded an accuracy of 99.9%, an MCC of 0.98, and a prediction time of 4316.57 microseconds (μs) on a Raspberry Pi device. For the NF-BoT-IoT-v2 dataset, the combination of ANOVA and Chi2 achieved 100% accuracy, an MCC of 0.99, and a prediction time of 4385.69&#xa0;μs on the same device.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

New design strategies for IoT intrusion detection using boosting and feature selection

  • Abderahmane Hamdouchi,
  • Ali Idri

摘要

The adoption of Internet of Things (IoT) technologies has introduced significant security challenges due to the diversity and number of connected devices. Although machine learning (ML) is widely used in intrusion detection systems (IDSs), it does not always ensure a good trade-off between bias and variance. This study presents a large-scale evaluation of ensemble learning methods, focusing on boosting techniques while considering feature types and transformation strategies. Five boosting algorithms (adaptive boosting (ADB), gradient boosting machines (GBM), category boosting (CATB), light gradient boosting machine (LGBM), and extreme gradient boosting (XGB)) were evaluated using two transformation strategies and five filter-based feature selection (FS) methods: analysis of variance (ANOVA), Kendall’s tau, mutual information (MI), maximum relevance minimum redundancy (mRMR), and Chi-square (Chi2). Experiments were conducted on two large NetFlow-based IoT datasets (NF-ToN-IoT-v2 and NF-BoT-IoT-v2). Each model was assessed using Matthews correlation coefficient (MCC), Cohen’s kappa, F1-score, and accuracy, and ranked using the Scott–Knott statistical test and Borda count voting system. A total of 1160 models were trained using the Toubkal Supercomputer, whose high-performance computing (HPC) infrastructure enabled scalable parallel and distributed processing. The best results were obtained using XGB with 200 estimators and 18 selected features under the standardization transformation. For the NF-ToN-IoT-v2 dataset, the combination of ANOVA and mRMR yielded an accuracy of 99.9%, an MCC of 0.98, and a prediction time of 4316.57 microseconds (μs) on a Raspberry Pi device. For the NF-BoT-IoT-v2 dataset, the combination of ANOVA and Chi2 achieved 100% accuracy, an MCC of 0.99, and a prediction time of 4385.69 μs on the same device.