This paper explores trimmed factorial k-means (TFKM) in a clustering application to a cookie dataset. TFKM is a robust version of factorial k-means, where a robust covariance matrix input is used, and outliers in the identified reduced space are iteratively removed via a trimming procedure. The selected latent rank, number of clusters, and outlier proportion are those which maximize Hartigan’s statistic. The TFKM partition is thoroughly compared to two alternatives, like a robust tandem procedure and trimmed k-means, via a simulation study. An Internet cookie example shows that TFKM provides on analyzed data the most parsimonious and informative partition by cluster homogeneity.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Trimmed Factorial K-Means: A Clustering Application to a Cookie Dataset

  • Matteo Farnè,
  • Furio Camillo

摘要

This paper explores trimmed factorial k-means (TFKM) in a clustering application to a cookie dataset. TFKM is a robust version of factorial k-means, where a robust covariance matrix input is used, and outliers in the identified reduced space are iteratively removed via a trimming procedure. The selected latent rank, number of clusters, and outlier proportion are those which maximize Hartigan’s statistic. The TFKM partition is thoroughly compared to two alternatives, like a robust tandem procedure and trimmed k-means, via a simulation study. An Internet cookie example shows that TFKM provides on analyzed data the most parsimonious and informative partition by cluster homogeneity.