Anomaly detection is the problem of finding patterns in data that do not conform to expected behavior. These non-conforming patterns are often referred to as discordant observations, outliers, aberrations, exceptions, anomalies, contaminants, peculiarities, or surprises in different application fields. In this study, anomalies were detected for 12 German cities using two supervised learning algorithms (Naive Bayes and XGBoost), and two unsupervised learning algorithms (Local Outlier Factor and autoencoder). LOF identifies anomalies by finding points in the dataset with remarkably lower density in contrast to their neighbors. Naïve Bayes identifies anomalies by calculating the likelihood of data points given the feature distributions. XGBoost identifies anomalies by the model to classify data points as outlier or regular established on the feature set. LOF identified minimal anomalies, indicating consistent local densities; the percentage of anomalies for each feature was between 0.01 and 1.68%, whereas in Naive Bayes anomalies detected ranged from 3 to 13% for different features and XGBoost detected moderate anomalies ranging from 2 to 6% for each feature, highlighting varying distributions. Autoencoder results varied significantly, with anomalies detected ranging from 9 to 95% for the weather dataset’s features. In the irradiation dataset, we had about 0.01–29% revealing its sensitivity to complex patterns. These findings demonstrate the importance of selecting appropriate algorithms for various anomaly detection scenarios and therefore provide a robust framework for active environmental management.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Anomaly Detection in Weather and Irradiation Data: A Study of 12 German Cities

  • Rafiat Bamimore Akodu,
  • Taiwo Sanda,
  • Rasheed Oyewole,
  • Ikram Chairi,
  • Robert Basmadjian

摘要

Anomaly detection is the problem of finding patterns in data that do not conform to expected behavior. These non-conforming patterns are often referred to as discordant observations, outliers, aberrations, exceptions, anomalies, contaminants, peculiarities, or surprises in different application fields. In this study, anomalies were detected for 12 German cities using two supervised learning algorithms (Naive Bayes and XGBoost), and two unsupervised learning algorithms (Local Outlier Factor and autoencoder). LOF identifies anomalies by finding points in the dataset with remarkably lower density in contrast to their neighbors. Naïve Bayes identifies anomalies by calculating the likelihood of data points given the feature distributions. XGBoost identifies anomalies by the model to classify data points as outlier or regular established on the feature set. LOF identified minimal anomalies, indicating consistent local densities; the percentage of anomalies for each feature was between 0.01 and 1.68%, whereas in Naive Bayes anomalies detected ranged from 3 to 13% for different features and XGBoost detected moderate anomalies ranging from 2 to 6% for each feature, highlighting varying distributions. Autoencoder results varied significantly, with anomalies detected ranging from 9 to 95% for the weather dataset’s features. In the irradiation dataset, we had about 0.01–29% revealing its sensitivity to complex patterns. These findings demonstrate the importance of selecting appropriate algorithms for various anomaly detection scenarios and therefore provide a robust framework for active environmental management.