The high-dimensional data analysis and feature selection play a pivotal role in enhancing model interpretability, reducing computational complexity, and improving prediction performance. One significant issue lies in the scalability of the method to extremely high-dimensional datasets, where computational complexity becomes a bottleneck. The objective of the study on unsupervised feature selection for high-dimensional data using clustering and multi-objective optimization is to develop a novel approach for efficiently identifying relevant features in high-dimensional datasets without requiring labelled information. The application of sophisticated feature clustering algorithms within Kernel Multiresolution Graph-Based Clustering (KMRGC) is crucial for successful unsupervised feature selection in high-dimensional datasets. In the feature selection process employing a Multi-Objective Genetic Algorithm (MOGA), this study optimizes two crucial objectives for relevance and non-redundancy. The results show that the algorithm's performance, with an accuracy of 74%, is influenced by the dimensionality of the feature space it operates on, this implementation utilizes Python software. The future trajectory of this research lies in the continual refinement and extension of the proposed approach to address evolving challenges in high-dimensional data analysis across diverse domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Feature Selection for High-Dimensional Data Using Clustering and Multi-Objective Optimization

  • Suman Laha,
  • Utpal Roy

摘要

The high-dimensional data analysis and feature selection play a pivotal role in enhancing model interpretability, reducing computational complexity, and improving prediction performance. One significant issue lies in the scalability of the method to extremely high-dimensional datasets, where computational complexity becomes a bottleneck. The objective of the study on unsupervised feature selection for high-dimensional data using clustering and multi-objective optimization is to develop a novel approach for efficiently identifying relevant features in high-dimensional datasets without requiring labelled information. The application of sophisticated feature clustering algorithms within Kernel Multiresolution Graph-Based Clustering (KMRGC) is crucial for successful unsupervised feature selection in high-dimensional datasets. In the feature selection process employing a Multi-Objective Genetic Algorithm (MOGA), this study optimizes two crucial objectives for relevance and non-redundancy. The results show that the algorithm's performance, with an accuracy of 74%, is influenced by the dimensionality of the feature space it operates on, this implementation utilizes Python software. The future trajectory of this research lies in the continual refinement and extension of the proposed approach to address evolving challenges in high-dimensional data analysis across diverse domains.