A dynamic method for preparing microarray gene expression data in disease classification system
摘要
The rapid growth of microarray gene expression (MAGE) datasets enhances the design of computational models for disease diagnosis, prognosis, prediction, and associated applications. Molecular classification involves supervised pattern recognition and includes the development of new categorization and gene selection models. Processing MAGE data presents challenges related to high dimensionality, complexity, and the presence of noise. Existing MAGE analysis methods lack reliability and flexibility in pre-processing MAGE data for disease classification. A new dynamic pre-processing strategy is proposed to address enormous dimensionality. The aim is to employ a threshold value dynamically to exclude gene data that is redundant and noisy. The input MAGE data is first processed to identify and discard the low-expressed gene profiles using different filters. Furthermore, the filtered gene data is converted into a frequency domain using the Discrete Cosine Transform (DCT). The threshold value is dynamically calculated for each gene dataset to ensure accurate identification and filtering of noisy and redundant gene profiles. Finally, the Inverse DCT (IDCT) is applied to reconstruct the pre-processed gene profiles in the time domain. The algorithm’s effectiveness is assessed using public datasets. The simulation results show that the suggested MAGE data reduction strategy removes 13.79% more redundant genes and decreases processing time by 6.04%.
Graphical abstract