Continual learning (CL) focuses on enabling machine learning algorithms to learn from a series of tasks without forgetting previously acquired knowledge. The use of continual learning has not been widely explored in cybersecurity and network safety applications, partially due to the lack of proper datasets. Besides, the benchmark datasets used in CL methods are often relatively restrictive in terms of data distribution shift among the tasks. In this work, we present a CL benchmark framework to construct datasets for CL in cybersecurity applications. For the cybersecurity applications, the proposed framework can generate datasets for CL under distribution shifts in data inputs (e.g., features of internet traffic flow), distribution shifts in data output (e.g., intrusion types), and distribution shifts in both data inputs and outputs, respectively. Moreover, we propose several distance-based and model-based metrics to meticulously quantify the magnitude of distribution shift between datasets of the tasks. We elaborate the construction of benchmark datasets and evaluate the quality of the constructed datasets by applying several existing CL methods and investigating their performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Composition for Continual Learning in Application of Cyberattack Detection

  • Jiayi Lian,
  • Xueying Liu,
  • Kevin Choi,
  • Balaji Veeramani,
  • Sathvik Murli,
  • Alison Hu,
  • Laura Freeman,
  • Edward Bowen,
  • Xinwei Deng

摘要

Continual learning (CL) focuses on enabling machine learning algorithms to learn from a series of tasks without forgetting previously acquired knowledge. The use of continual learning has not been widely explored in cybersecurity and network safety applications, partially due to the lack of proper datasets. Besides, the benchmark datasets used in CL methods are often relatively restrictive in terms of data distribution shift among the tasks. In this work, we present a CL benchmark framework to construct datasets for CL in cybersecurity applications. For the cybersecurity applications, the proposed framework can generate datasets for CL under distribution shifts in data inputs (e.g., features of internet traffic flow), distribution shifts in data output (e.g., intrusion types), and distribution shifts in both data inputs and outputs, respectively. Moreover, we propose several distance-based and model-based metrics to meticulously quantify the magnitude of distribution shift between datasets of the tasks. We elaborate the construction of benchmark datasets and evaluate the quality of the constructed datasets by applying several existing CL methods and investigating their performance.