Classification of Cancer Types Based on RNA HI-SEQ Data Using Dimensionality Reduction
摘要
Many attempts have been made to enhance the accuracy of cancer classification through the use of gene expression data. However, the use of extensive data increases the potential for data overfitting, highlighting the need for more efficient approaches. This paper aims to find a more convenient method to classify the genome sequencing datasets. We have used three types of dimensionality reduction to reduce the feature number to make the dataset effective to use. We used the Pancan Hi-sequence dataset, which included 801 cases and 20531 genes for each case, resulting in fewer principle components using PCA, t-SNE, and UMAP. Initially, the data was transformed from a high-dimensional feature space to a lower-dimensional one through the reduction processes. Subsequently, the reduced features were assessed using both linear and non-linear SVM models with various kernels, resulting in an almost 99% accuracy rate.