The rapid growth of microarray gene expression (MAGE) databases allows computer models for disease diagnosis, prognosis, prediction, etc. Soft computing, data mining, machine learning, and pattern recognition have helped researchers develop computationally effective models to identify new disease classes gain insight into disease mechanisms, and develop diagnostic or therapeutic targets through molecular classification. Molecular classification is a supervised pattern recognition problem that challenges the development of novel classification and gene selection models by addressing high dataset dimensionality, overfitting due to imbalanced data elements in different classes, and computational complexity. This paper focused on the high dimensionality problem by proposing the dynamic pre-processing algorithm. The purpose of this approach is to use a dynamically determined threshold value to eliminate redundant and noisy gene data. The complex MAGE data is converted into the simple using the proposed pre-processing algorithm. The core part of this model is converting the input MAGE data using Discrete Cosine Transform (DCT), applying dynamic thresholding, and then performing inverse DCT. Using publicly accessible datasets, the suggested algorithm’s effectiveness is assessed. The simulation results revealed the proposed MAGE data reduction algorithm discards the redundant genes by 13.79% more and minimizes the processing time by 6.04% compared to existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic Approach for Pre-processing of Microarray Gene Expression Data

  • Hemant B. Mahajan,
  • K. T. V. Reddy

摘要

The rapid growth of microarray gene expression (MAGE) databases allows computer models for disease diagnosis, prognosis, prediction, etc. Soft computing, data mining, machine learning, and pattern recognition have helped researchers develop computationally effective models to identify new disease classes gain insight into disease mechanisms, and develop diagnostic or therapeutic targets through molecular classification. Molecular classification is a supervised pattern recognition problem that challenges the development of novel classification and gene selection models by addressing high dataset dimensionality, overfitting due to imbalanced data elements in different classes, and computational complexity. This paper focused on the high dimensionality problem by proposing the dynamic pre-processing algorithm. The purpose of this approach is to use a dynamically determined threshold value to eliminate redundant and noisy gene data. The complex MAGE data is converted into the simple using the proposed pre-processing algorithm. The core part of this model is converting the input MAGE data using Discrete Cosine Transform (DCT), applying dynamic thresholding, and then performing inverse DCT. Using publicly accessible datasets, the suggested algorithm’s effectiveness is assessed. The simulation results revealed the proposed MAGE data reduction algorithm discards the redundant genes by 13.79% more and minimizes the processing time by 6.04% compared to existing methods.