<p>One of the most popular partitioning cluster algorithms is k-means, which is only applicable to numerical data. An extension to mixed-type data containing numerical and categorical variables is the k-prototypes algorithm. Due to its iterative structure, the algorithm may only converges to a local minimum rather than a global minimum. Therefore, just like the solution of the original k-means, the resulting cluster partition suffers from the initialization. In general, there are two ways of achieving an improvement of the random-based initialization of the algorithm: One possibility is to determine concrete initial cluster centers, and the other strategy is to repeat the algorithm with different randomly chosen initial centers. In this work, algorithm initializations of both options are analyzed and evaluated comparatively in a benchmark study. Therefore, selected initialization strategies of the k-means algorithm are transformed to the application on mixed-type data. For the simulation study, several data sets are artificially generated and cluster partitions are determined by using the competing initialization strategies. It is shown that an improvement of the cluster algorithm’s target criterion can be achieved as well as the ability to identify appropriate groups, even with manageable time expenditure.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Initialization strategies for clustering mixed-type data with the k-prototypes algorithm

  • Rabea Aschenbruck,
  • Gero Szepannek,
  • Adalbert F. X. Wilhelm

摘要

One of the most popular partitioning cluster algorithms is k-means, which is only applicable to numerical data. An extension to mixed-type data containing numerical and categorical variables is the k-prototypes algorithm. Due to its iterative structure, the algorithm may only converges to a local minimum rather than a global minimum. Therefore, just like the solution of the original k-means, the resulting cluster partition suffers from the initialization. In general, there are two ways of achieving an improvement of the random-based initialization of the algorithm: One possibility is to determine concrete initial cluster centers, and the other strategy is to repeat the algorithm with different randomly chosen initial centers. In this work, algorithm initializations of both options are analyzed and evaluated comparatively in a benchmark study. Therefore, selected initialization strategies of the k-means algorithm are transformed to the application on mixed-type data. For the simulation study, several data sets are artificially generated and cluster partitions are determined by using the competing initialization strategies. It is shown that an improvement of the cluster algorithm’s target criterion can be achieved as well as the ability to identify appropriate groups, even with manageable time expenditure.