Principal Components-Based Classification Using a Linear Discriminant Analyzer and Enhancement of its Prediction Accuracy
摘要
Principal Component Analysis (PCA) is a preferred technique for dimensionality reduction of complex multi-dimensional datasets. Though PCA generates n principal components (PCs) for an n-dimensional dataset, we retain only the p (<< n) high variance-PCs (variance ≥ 1.0) that are considered to maximally capture variations in the feature values in the original dataset. We show that the prediction accuracy of principal components-based classification could be almost equal to or even larger than that of classifiers modeled using all the features of the original dataset. To improve the prediction accuracy of PCs-based classification, we propose a hybrid training dataset comprising of the p retained PCs and a raw feature with a relatively moderate-high variance in the original dataset. We show that the prediction accuracy of classifiers modeled using such a hybrid training dataset could match or exceed to that of classifiers modeled using all the features in the original dataset.