错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Semi-supervised Framework for Anomaly Detection and Data Labeling for Industrial Control Systems

  • Jiyan Salim Mahmud,
  • Ermiyas Birihanu,
  • Imre Lendak

摘要

To ensure uninterrupted service delivery in critical sectors like electricity, water, and oil, safeguarding information systems against anomalies is imperative. Detecting anomalies within Industrial Control Systems (ICSs) is vital, but it’s challenging without a comprehensive understanding of their causes. This necessitates a well-annotated dataset encompassing diverse anomaly types, often dependent on domain experts. Unfortunately, such datasets are scarce. To address this challenge, this study introduces a specialized framework for unsupervised anomaly detection and anomaly categorization within data collected from monitoring ICSs. The framework was validated using data from a Secure Water Treatment (SWaT) testbed, where multiple cyberattacks were intentionally introduced. An Isolation Forest model was utilized, achieving 77% accuracy in anomaly identification. These anomalies were then isolated from normal samples, and a K-means clustering model categorized similar attacks and labeled anomaly clusters. The most suitable supervised model for the data was determined through experimentation with various classifiers, including SVM, Random Forest, Decision Tree, KNearest Neighbor, and AdaBoost. Remarkably, K-Nearest Neighbor (KNN) outperformed all, achieving 98% accuracy. This framework automates anomaly detection, categorization and data labeling, elevating data quality and accuracy in ICS anomaly detection while reducing the need for manual expert intervention and addressing the challenge of limited well-annotated datasets and improving the overall security of vital infrastructure sectors.