Dealing with high-dimensional datasets is challenging nowadays due to computational complexity, the curse of dimensionality, and model overfitting. It becomes necessary to reduce the dimension of the dataset for a better understanding of inherent information. Feature selection techniques are widely utilized to rank features based on their importance and accordingly reduce the dimension of the original datasets with respect to this ranking. Existing feature selection methods are mainly developed for specific downstream tasks and show several drawbacks e.g., not considering the inherent associations and their importance. In most of the cases, the methods are computationally expensive as well. In order to address such drawbacks, the present study aims to propose a feature selection method, called NeuroDAVIS-FS, which performs in an unsupervised learning setup without assuming any prior data distribution. Initially, it considers training using the model NeuroDAVIS, developed earlier for data visualization, and selects features according to the trained model. The efficacy of the proposed NeuroDAVIS-FS has been demonstrated on various datasets from different domains and found to be effective in comparison with state-of-the-art feature selection methods. In addition, two case studies on image and biological datasets with a very low sample-feature ratio, have been executed and found to be effective for relevant feature selection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

NeuroDAVIS-FS: Feature Selection Through Visualization Using NeuroDAVIS

  • Chayan Maitra,
  • Anwesha Sengupta,
  • Rajat K. De

摘要

Dealing with high-dimensional datasets is challenging nowadays due to computational complexity, the curse of dimensionality, and model overfitting. It becomes necessary to reduce the dimension of the dataset for a better understanding of inherent information. Feature selection techniques are widely utilized to rank features based on their importance and accordingly reduce the dimension of the original datasets with respect to this ranking. Existing feature selection methods are mainly developed for specific downstream tasks and show several drawbacks e.g., not considering the inherent associations and their importance. In most of the cases, the methods are computationally expensive as well. In order to address such drawbacks, the present study aims to propose a feature selection method, called NeuroDAVIS-FS, which performs in an unsupervised learning setup without assuming any prior data distribution. Initially, it considers training using the model NeuroDAVIS, developed earlier for data visualization, and selects features according to the trained model. The efficacy of the proposed NeuroDAVIS-FS has been demonstrated on various datasets from different domains and found to be effective in comparison with state-of-the-art feature selection methods. In addition, two case studies on image and biological datasets with a very low sample-feature ratio, have been executed and found to be effective for relevant feature selection.