错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Role of Pre-processing in Gene Selection Using DNA Microarray Gene Expression Data

  • Tanusri Ghosh,
  • Sriyankar Acharyya

摘要

Precision medicine is a boon to the medical field recently for early disease detection, monitoring disease progression, and developing new drugs. To successfully deploy precision medicine, biomarker identification is the first step and DNA microarray technology acts as a powerful tool in recent decades. Using DNA microarray data, one can analyze the tens of thousands of genes simultaneously, but it also has some limitations. The data is quite noisy and contains irrelevant genes. Furthermore, it has a dimension imbalance problem which affects the overall performance of the gene selection process or disease classification accuracy. This paper has emphasized on the role of pre-processing approach using some widely used statistical methods to overcome these drawbacks. To show the importance of pre-processing, here, there are two different approaches: gene selection without pre-processing and with pre-processing on two real-life datasets such as Breast Cancer and Leukemia. Gene selection is viewed here as an optimization problem, and the optimization is done using the proposed PSO-SVM gene selection model. It can be observed that the second approach (with pre-processing) gives better results as compared to the former. After comparing the performance of the pre-processing methods, based on gene selection (maximum accuracy was achieved by choosing minimum genes) it can be inferred that SNR and Fisher Score are competitive and better than others.