<p>Clustering mixed-type data containing numeric, nominal, and ordinal attributes poses challenges, such as distances that do not adapt during the clustering process and imbalances in attribute type proportions that can lead to one type dominating the results. This study proposes an Entropy-Weighted Learnable Mixed-type Data Distance (EW-LMD) algorithm that combines three components: (a) a mixed distance function that distinguishes nominal and ordinal treatments based on the hierarchical structure of values; (b) a learnable distance mechanism that iteratively updates distances from the probability of data co-occurrence within a cluster; and (c) an entropy-based attribute weighting that prioritizes informative attributes to suppress data type bias. The algorithm is available in two versions: a manual version that requires configuring the distance constant, learning rate, number of bins, and number of clusters; and an autonomous version that estimates these parameters based on dataset characteristics. Evaluation on fifteen datasets from various domains demonstrates the superiority of EW-LMD over K-Prototypes, KAMILA, MFCM, and other comparable methods. The autonomous version outperformed in entropy reduction across nine datasets, cluster utility across ten datasets, and external metrics ACC, ARI, and NMI across eight to nine datasets. Furthermore, the autonomous version was more consistent across datasets with more stable medians and quartiles. However, the autonomous version of EW-LMD required 35% more computation time than the manual version due to early parameter estimation. The main contribution of this research is that the algorithm overcomes the limitations of static distance by leveraging adaptive distance learning and attribute-level entropy weighting, resulting in more structurally homogeneous clusters and better alignment with the ground truth.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Entropy-weighted learnable distance clustering for mixed-type data

  • Fitri Nuraeni,
  • Abdul Syukur,
  • Aris Marjuni,
  • Nova Rijati,
  • Dede Kurniadi

摘要

Clustering mixed-type data containing numeric, nominal, and ordinal attributes poses challenges, such as distances that do not adapt during the clustering process and imbalances in attribute type proportions that can lead to one type dominating the results. This study proposes an Entropy-Weighted Learnable Mixed-type Data Distance (EW-LMD) algorithm that combines three components: (a) a mixed distance function that distinguishes nominal and ordinal treatments based on the hierarchical structure of values; (b) a learnable distance mechanism that iteratively updates distances from the probability of data co-occurrence within a cluster; and (c) an entropy-based attribute weighting that prioritizes informative attributes to suppress data type bias. The algorithm is available in two versions: a manual version that requires configuring the distance constant, learning rate, number of bins, and number of clusters; and an autonomous version that estimates these parameters based on dataset characteristics. Evaluation on fifteen datasets from various domains demonstrates the superiority of EW-LMD over K-Prototypes, KAMILA, MFCM, and other comparable methods. The autonomous version outperformed in entropy reduction across nine datasets, cluster utility across ten datasets, and external metrics ACC, ARI, and NMI across eight to nine datasets. Furthermore, the autonomous version was more consistent across datasets with more stable medians and quartiles. However, the autonomous version of EW-LMD required 35% more computation time than the manual version due to early parameter estimation. The main contribution of this research is that the algorithm overcomes the limitations of static distance by leveraging adaptive distance learning and attribute-level entropy weighting, resulting in more structurally homogeneous clusters and better alignment with the ground truth.