<p>This study introduces and evaluates the Euclidean Distance-Based Outlier Detection Method (EDBODM) as a new approach for detecting outliers in grouped multivariate data. The evaluation considers both equal and unequal covariance scenarios, analyzing uncontaminated data as well as data with 6% contamination across various multivariate distributions. The results consistently highlight the superior performance of EDBODM over established methods by Hardin and Rocke (Comput Stat Data Anal 44:625–638, 2004) and Caroni and Billor (J Appl Stat 34:1241–1250, 2007), particularly in terms of specificity. It is strong enough to find outliers accurately, no matter what the underlying multivariate distribution or covariance structure is like in clean data. Additionally, EDBODM is more accurate, precise, and recalling when dealing with contaminated data, such as when the data has multivariate normal or skew-t distributions. Its balanced detection approach minimizes false positives while ensuring reliable identification of true outliers. These results show that EDBODM is a powerful tool for finding strong outliers in grouped multivariate data that can be used in a wide range of situations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EDBODM: A Euclidean Distance-Based Method for Detecting Outliers in Grouped Multivariate Data

  • Suthat Phuttisen,
  • Wuttichai Srisodaphol

摘要

This study introduces and evaluates the Euclidean Distance-Based Outlier Detection Method (EDBODM) as a new approach for detecting outliers in grouped multivariate data. The evaluation considers both equal and unequal covariance scenarios, analyzing uncontaminated data as well as data with 6% contamination across various multivariate distributions. The results consistently highlight the superior performance of EDBODM over established methods by Hardin and Rocke (Comput Stat Data Anal 44:625–638, 2004) and Caroni and Billor (J Appl Stat 34:1241–1250, 2007), particularly in terms of specificity. It is strong enough to find outliers accurately, no matter what the underlying multivariate distribution or covariance structure is like in clean data. Additionally, EDBODM is more accurate, precise, and recalling when dealing with contaminated data, such as when the data has multivariate normal or skew-t distributions. Its balanced detection approach minimizes false positives while ensuring reliable identification of true outliers. These results show that EDBODM is a powerful tool for finding strong outliers in grouped multivariate data that can be used in a wide range of situations.