<p>Outliers are known to be detrimental to widely used clustering techniques. Robust clustering alternatives have been introduced to better resist outlying observations. Among these, robust clustering methods based on trimming have proven effective by allowing the removal of a fraction of observations where outliers are likely to be found, with TCLUST being one of the most popular for handling elliptically contoured clusters. The algorithm for applying TCLUST can be seen as an extension of the concentration steps used in the fast-MCD algorithm for computing the Minimum Covariance Determinant. However, obtaining good initializations for these concentration steps in TCLUST is more complex than in MCD. This initialization task is particularly challenging unless both the number of clusters and the dimensionality are small. To address this, a new ensemble initialization procedure for TCLUST will be presented, which takes advantage of partially correct information from all iterated random initializations rather than focusing solely on the best individual one found. Initial experiments suggest that this methodology could improve the computational performance of the standard TCLUST algorithm. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving the computational performance of TCLUST through ensemble initialization

  • Pedro C. Álvarez-Esteban,
  • Luis A. García-Escudero,
  • Agustín Mayo-Iscar,
  • Javier Crespo-Guerrero

摘要

Outliers are known to be detrimental to widely used clustering techniques. Robust clustering alternatives have been introduced to better resist outlying observations. Among these, robust clustering methods based on trimming have proven effective by allowing the removal of a fraction of observations where outliers are likely to be found, with TCLUST being one of the most popular for handling elliptically contoured clusters. The algorithm for applying TCLUST can be seen as an extension of the concentration steps used in the fast-MCD algorithm for computing the Minimum Covariance Determinant. However, obtaining good initializations for these concentration steps in TCLUST is more complex than in MCD. This initialization task is particularly challenging unless both the number of clusters and the dimensionality are small. To address this, a new ensemble initialization procedure for TCLUST will be presented, which takes advantage of partially correct information from all iterated random initializations rather than focusing solely on the best individual one found. Initial experiments suggest that this methodology could improve the computational performance of the standard TCLUST algorithm.