错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting Anomalies in Data Using Z-numbers

  • Alekperov Ramiz Balashirin,
  • Sufanzade Tural Togrul

摘要

The study of issues related to the detection of anomalies in time series in various fields of activity is very relevant. There are several methods for detecting anomalies in time series. However, there are some problems with detecting anomalies in the data, and understanding them is important for developing more effective approaches.: - The definition of an anomaly is mostly subjective and depends on the context. - The lack of clear definitions of anomalies makes it difficult to develop effective models.- In real-world scenarios, it is often difficult to obtain training data with enough explicit examples of anomalies, which makes it difficult to train models. Work to eliminate or mitigate these deficiencies is an active area of research in the field of detecting anomalies in data. On the other hand, with the development of data collection technologies, the volume of time series increases significantly over time. Automated anomaly detection methods are becoming critically important when processing and analyzing such volumes of information. In the article, to identify anomalies in data, it is proposed to use Z–numbers that take into account the peculiarities of human thinking, which is based on the experience and intuition of experts who are more qualified to identify anomalies in data and reason in the following context, for example, if the value of x is 56, it is suspicious and there is great confidence in this. The idea is that the assessment of anomalies is carried out taking into account the deviation of real data (which are converted into Z-numbers in advance) from possible options presented as Z–numbers and after estimating the distance between them, their degree of normality is ranked, and determined with a certain confidence in the form of finite Z–numbers. The results obtained using this approach can be useful to specialists from various companies who need to solve data-cleaning issues.