Time-series anomaly detection (TAD) has a pivotal role across various domains ranging from manufacturing to health care monitoring. Numerous machine learning solutions have been proposed for TAD, with varying levels of complexity. However, most solutions benchmark their performance using misleading evaluation metrics which hinder reliable comparative analysis and the development of truly robust TAD methods. In the present work, we disentangle how performance evaluation can be unreliable due to several factors: suboptimal scoring functions, thresholding functions that assume access to all test labels, prediction modification based on test labels, lack of benchmarking against trivial baselines, and finally, problematic datasets. In this paper, we endeavor to address these issues by introducing a comprehensive TAD evaluation framework which includes: state-of-the-art deep-learning (DL) and traditional machine learning (ML) TAD algorithms; TAD baselines; an extensive set of scoring, thresholding and evaluation functions. Our rigorous analysis shows that: (i) TAD baselines and simple ML algorithms achieve performance often on par with advanced SOTA DL solutions. (ii) Scoring and thresholding function selection can greatly impact the anomaly prediction performance. (iii) Evaluation metrics used in the field, mostly focused on post-thresholding output, are worryingly inconsistent and can generate starkly overestimated predictions. We advocate instead for a more widespread use of pre-thresholding metrics and for post-thresholding metrics that closely correlate to the former. Our code is available at https://github.com/intellabs/tsad-ef .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Robust Framework for Evaluation of Unsupervised Time-Series Anomaly Detection

  • Onat Gungor,
  • Amanda Rios,
  • Priyanka Mudgal,
  • Nilesh Ahuja,
  • Tajana Rosing

摘要

Time-series anomaly detection (TAD) has a pivotal role across various domains ranging from manufacturing to health care monitoring. Numerous machine learning solutions have been proposed for TAD, with varying levels of complexity. However, most solutions benchmark their performance using misleading evaluation metrics which hinder reliable comparative analysis and the development of truly robust TAD methods. In the present work, we disentangle how performance evaluation can be unreliable due to several factors: suboptimal scoring functions, thresholding functions that assume access to all test labels, prediction modification based on test labels, lack of benchmarking against trivial baselines, and finally, problematic datasets. In this paper, we endeavor to address these issues by introducing a comprehensive TAD evaluation framework which includes: state-of-the-art deep-learning (DL) and traditional machine learning (ML) TAD algorithms; TAD baselines; an extensive set of scoring, thresholding and evaluation functions. Our rigorous analysis shows that: (i) TAD baselines and simple ML algorithms achieve performance often on par with advanced SOTA DL solutions. (ii) Scoring and thresholding function selection can greatly impact the anomaly prediction performance. (iii) Evaluation metrics used in the field, mostly focused on post-thresholding output, are worryingly inconsistent and can generate starkly overestimated predictions. We advocate instead for a more widespread use of pre-thresholding metrics and for post-thresholding metrics that closely correlate to the former. Our code is available at https://github.com/intellabs/tsad-ef .