<p>An outlier is a data point that is significantly distant from other data points in a dataset, which can bias the estimator or cause estimation algorithms to fail to converge. Thus, it is recommended that outliers be detected and handled properly before the analysis. In this paper, we study a procedure for detecting outliers when the data of interest are histograms. We propose a one-class support vector machine (one-SVM) based on the Wasserstein distance, where, more specifically, an exponential kernel with the Wasserstein distance is used. Through numerical experiments, we show that the one-SVM with the Wasserstein distance outperforms those based on other histogram distances, namely the Euclidean, Kolmogorov-Smirnov, and total variation distances. We apply the one-SVM procedure to yearly histogram data that record daily temperatures from January 1, 1960, to December 31, 2023, for five major cities—Seoul, Busan, Daegu, Gwangju, Daejeon—in Korea. In the analysis, we aim to detect outlying years in which the distribution of average daily temperatures deviates from the typical pattern.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Outlier Detection for Histogram-Valued Data using Wasserstein Distance based One Class SVM

  • Seul Lee,
  • Yangha Chung,
  • Soohyun Ahn,
  • Johan Lim

摘要

An outlier is a data point that is significantly distant from other data points in a dataset, which can bias the estimator or cause estimation algorithms to fail to converge. Thus, it is recommended that outliers be detected and handled properly before the analysis. In this paper, we study a procedure for detecting outliers when the data of interest are histograms. We propose a one-class support vector machine (one-SVM) based on the Wasserstein distance, where, more specifically, an exponential kernel with the Wasserstein distance is used. Through numerical experiments, we show that the one-SVM with the Wasserstein distance outperforms those based on other histogram distances, namely the Euclidean, Kolmogorov-Smirnov, and total variation distances. We apply the one-SVM procedure to yearly histogram data that record daily temperatures from January 1, 1960, to December 31, 2023, for five major cities—Seoul, Busan, Daegu, Gwangju, Daejeon—in Korea. In the analysis, we aim to detect outlying years in which the distribution of average daily temperatures deviates from the typical pattern.