Anomaly detection is crucial across various domains such as fraud detection, health care, network security, and monitoring industrial control systems. A significant challenge in this field is evaluating the performance of unsupervised anomaly detection algorithms due to the absence of labeled data in real-world applications. This study explores the application of Silhouette Coefficient to assess the outcomes of different unsupervised anomaly detection algorithms. We conducted multiple experiments across two public industrial control system datasets, using three unsupervised models: One-Class SVM, Isolation Forest, and Local Outlier Factor. By comparing the results of Silhouette Coefficient and traditional supervised metrics F1-score and Area Under the Receiver Operating Characteristic Curve, we observed a positive correlation between these metrics and Silhouette Coefficient. This finding indicates that higher F1-scores tend to be associated with better-defined clusters, suggesting that Silhouette Coefficient is a reliable metric for evaluating the effectiveness of unsupervised anomaly detection algorithms without labeled data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Unsupervised Anomaly Detection in Industrial Systems Using Silhouette Coefficient

  • Jiyan Salim Mahmud,
  • Ali Youssef,
  • Imre Lendák

摘要

Anomaly detection is crucial across various domains such as fraud detection, health care, network security, and monitoring industrial control systems. A significant challenge in this field is evaluating the performance of unsupervised anomaly detection algorithms due to the absence of labeled data in real-world applications. This study explores the application of Silhouette Coefficient to assess the outcomes of different unsupervised anomaly detection algorithms. We conducted multiple experiments across two public industrial control system datasets, using three unsupervised models: One-Class SVM, Isolation Forest, and Local Outlier Factor. By comparing the results of Silhouette Coefficient and traditional supervised metrics F1-score and Area Under the Receiver Operating Characteristic Curve, we observed a positive correlation between these metrics and Silhouette Coefficient. This finding indicates that higher F1-scores tend to be associated with better-defined clusters, suggesting that Silhouette Coefficient is a reliable metric for evaluating the effectiveness of unsupervised anomaly detection algorithms without labeled data.