Intrusion Detection Systems (IDS) are critical for securing future communication systems, yet existing evaluation methods lack standardization, resulting in incomplete and unreliable assessments. In particular, the evaluation of machine-learning (ML) based IDS often boils down to the demonstration that the IDS performs well on a given dataset, regardless of the dataset quality. Prior evaluation approaches lack formalization and disregard ML best practices. This paper addresses this challenge by presenting FREIDA, a concrete tool for ensuring completeness, reliability, and reproducibility of ML-based IDS evaluations. This tool emphasizes the relationship between evaluation choices and data selection, requiring the generation of purpose-specific datasets. In this paper, we present and provide a Python implementation of our evaluation tool, that is used to evaluate multiple models in a variety of settings. This research represents a crucial stride toward standardizing IDS evaluation methods. Indeed, the proposed tool facilitates the systematic evaluation of IDS by producing a generated configuration file, thereby enhancing the reproducibility of each evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FREIDA: A Concrete Tool for Reproducible Evaluation of IDS Using a Data-Driven Approach

  • Solayman Ayoubi,
  • Gregory Blanc,
  • Houda Jmila,
  • Sébastien Tixeuil

摘要

Intrusion Detection Systems (IDS) are critical for securing future communication systems, yet existing evaluation methods lack standardization, resulting in incomplete and unreliable assessments. In particular, the evaluation of machine-learning (ML) based IDS often boils down to the demonstration that the IDS performs well on a given dataset, regardless of the dataset quality. Prior evaluation approaches lack formalization and disregard ML best practices. This paper addresses this challenge by presenting FREIDA, a concrete tool for ensuring completeness, reliability, and reproducibility of ML-based IDS evaluations. This tool emphasizes the relationship between evaluation choices and data selection, requiring the generation of purpose-specific datasets. In this paper, we present and provide a Python implementation of our evaluation tool, that is used to evaluate multiple models in a variety of settings. This research represents a crucial stride toward standardizing IDS evaluation methods. Indeed, the proposed tool facilitates the systematic evaluation of IDS by producing a generated configuration file, thereby enhancing the reproducibility of each evaluation.