This study compared several dimensionality reduction methods to analyze complex, high-dimensional data in adolescent development. Initially, we use different encoding methods, such as binary, ordinal, and one-hot encoding methods, to ensure the dataset's consistency. Then, to make our dataset less complicated, we used kernel PCA, t-Distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), Locally Linear Embedding (LLE), and Isometric Mapping (Isomap). We used the K-means clustering algorithm to find important patterns and then used different validation metrics, such as the Silhouette Score, Davies-Bouldin Score, and Calinski-Harabasz Score, to compare how well different dimensionality reduction techniques worked. Our results showed that that Isomap and UMAP consistently show the best performance across all metrics, with Isomap leading slightly. For instance, Isomap’s SS is 11.5% higher than UMAP's, and it also outperforms UMAP in DBS and CHS by 17.1 and 4.3%, respectively. Kernel PCA performs decently but falls short, with Isomap surpassing it by 27.0% in compactness and 67.2% in cluster separation. LLE and t-SNE are the weakest, with Isomap's SS being 11.9% higher than LLE’s and 22.5% higher than t-SNE's, and its CHS significantly outperforming both, by 188.2 and 73.0%, respectively. This study provides valuable insights into adolescent development, offering a clearer understanding of patterns and relationships, and suggesting that Isomap is particularly effective in capturing hidden features within the data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Advanced Dimensionality Reduction Techniques for Effective Clustering of High-Dimensional Adolescent Development Data

  • Khoula Said Al. Abri,
  • Manjit Singh Sidhu

摘要

This study compared several dimensionality reduction methods to analyze complex, high-dimensional data in adolescent development. Initially, we use different encoding methods, such as binary, ordinal, and one-hot encoding methods, to ensure the dataset's consistency. Then, to make our dataset less complicated, we used kernel PCA, t-Distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), Locally Linear Embedding (LLE), and Isometric Mapping (Isomap). We used the K-means clustering algorithm to find important patterns and then used different validation metrics, such as the Silhouette Score, Davies-Bouldin Score, and Calinski-Harabasz Score, to compare how well different dimensionality reduction techniques worked. Our results showed that that Isomap and UMAP consistently show the best performance across all metrics, with Isomap leading slightly. For instance, Isomap’s SS is 11.5% higher than UMAP's, and it also outperforms UMAP in DBS and CHS by 17.1 and 4.3%, respectively. Kernel PCA performs decently but falls short, with Isomap surpassing it by 27.0% in compactness and 67.2% in cluster separation. LLE and t-SNE are the weakest, with Isomap's SS being 11.9% higher than LLE’s and 22.5% higher than t-SNE's, and its CHS significantly outperforming both, by 188.2 and 73.0%, respectively. This study provides valuable insights into adolescent development, offering a clearer understanding of patterns and relationships, and suggesting that Isomap is particularly effective in capturing hidden features within the data.