Improved Clustering-Based Feature Selection Using Feature Extraction Based on Principal Component Analysis
摘要
Recently, big data on the phenomenon observed was gained easily. Nevertheless, this data certainly has a high dimensional due to the enormous features involved. Consequently, the accuracy of the analysis is low because of the irrelevant features involved. Besides, the high time complexity will occur. Therefore, several approaches are proposed to overcome it. One of them is feature selection (FS). This study proposed a hybrid model incorporating feature extraction using PCA into clustering for FS termed PCA-clustering, where five algorithms were proposed, namely PCA-KM, PCA-AL, PCA-CL, PCA-WL, and PCA-SL. The UCI dataset was utilized in this study, and three simulation schemes were conducted namely the 75%, 50%, and 25% schemes. The validation used was the goodness-of-fit of proximity matrix (GoFPM), the classification accuracy, the clustering results, and the time complexity. The existing FS algorithms and the performance of non-FS were utilized to evaluate the performance of the proposed algorithms. Generally, PCA-KM outperformed other proposed algorithms. Besides, PCA-KM had better performance in GoFPM, classification, and clustering than the existing algorithms in the particular conditions, especially at the 75% schemes. Meanwhile, in the 50% and 25% schemes, the proposed algorithms were under several existing algorithms. In terms of computational efficiency, the proposed algorithms were efficient, with a time complexity categorized as fair and a low computational time. Those showed that the proposed algorithms are viable for big data analysis.