<p>Extreme values or outliers are patterns in the data, which do not conform to a well-defined concept of normal behavior and usually appear to be produced by a different mechanism from the rest of the data. Determining an extreme value is a challenge for both humans and computers. In this paper, we propose a new approach to detect extreme values in univariate time series based on bar visibility. More specifically, the proposed approach is based on the calculation of the angles formed between bars of time series bar charts, supposing that vertical bars represent individual values at distinct points in time. Based on this approach, we propose three algorithms to detect extreme values. The three suggested solutions' main premise is that the extreme values appear to be "seen" with smaller angles from bars indicating the time series' normal values. The accuracy of the proposed approach is confirmed by experiments on both synthetic and real datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting extreme values in time series based on bar visibility

  • Maria Katsouda,
  • Basilis Boutsinas

摘要

Extreme values or outliers are patterns in the data, which do not conform to a well-defined concept of normal behavior and usually appear to be produced by a different mechanism from the rest of the data. Determining an extreme value is a challenge for both humans and computers. In this paper, we propose a new approach to detect extreme values in univariate time series based on bar visibility. More specifically, the proposed approach is based on the calculation of the angles formed between bars of time series bar charts, supposing that vertical bars represent individual values at distinct points in time. Based on this approach, we propose three algorithms to detect extreme values. The three suggested solutions' main premise is that the extreme values appear to be "seen" with smaller angles from bars indicating the time series' normal values. The accuracy of the proposed approach is confirmed by experiments on both synthetic and real datasets.