错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unleashing the power of machine learning in cancer analysis: a novel gene selection and classifier ensemble strategy

  • Jogeswar Tripathy,
  • Rasmita Dash,
  • Binod Kumar Pattanayak

摘要

Purpose

Globally, cancer is the second largest cause of mortality. For the improvement of cancer diagnosis, gene expression data plays a significant role. Cancer detection using traditional approaches is too complex and time-consuming. As an alternative, machine learning techniques are the better option in terms of computational cost for this critical task. However, the analysis of these data is too complex as the raw data is huge, noisy, and contains redundant genes. Thus, effective preprocessing and a robust data classification strategy are required to be designed for cancer data analysis.

Methods

This research is based on designing a novel feature selection and data classification technique. The proposed technique begins by removing noisy genes from the data using gene selection techniques. For this, a combinational approach is followed, in which a pool of ordered four gene ranking approaches are adopted, ordered pipelines are built, and significant genes are extracted. Furthermore, not to bias with single classifier performance, a classifier ensemble is prepared considering five simple and improved extreme learning machine (ELM) models with the soft voting scheme for efficient data classification purposes. This experiment is conducted over seven microarray datasets.

Results

After feature selection, ten frequently appearing features out of 16 pipelines are extracted for each dataset separately. Then, the proposed ensemble is compared with each individual classifier and the outcome is presented using four performance metrics. Overall, for all datasets, the performance of the ensemble is better over more than 74% of outcomes.

Conclusion

The concluding remark highlights the performance of the proposed design over a few considered state-of-the-art approaches.