Advances in high-throughput technologies have accelerated omics data research. The recent explosion of omics, namely transcriptomic data, has opened new opportunities for the discovery of novel biomarkers with potential to be incorporated into clinical practice. However, due to their extreme complexity, gaining useful insights is particularly challenging. Hence, the application of machine learning techniques on transcriptomic data emerges as a highly promising area for the discovery of new biomarkers. For exploring the potential of these techniques, this paper proposes a novel approach to process gene expression data with the aim of finding candidate gene signatures. Our methodology consists of an ensemble feature selection strategy based on the Boruta, SVM-RFE and LASSO methods, complemented by a second feature selection based on the gene importance calculated by the Random Forest, XGBoost, Support Vector Machine, Logistic Regression and AdaBoost methods. Performing simulations with a dataset of atopic dermatitis patients, our proposal resulted in an 8-gene signature with high AUC (0.839) and accuracy (0.8462) values.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Omics Data to Candidate Genes: An Innovative Machine Learning Approach for Biomarker Identification

  • Ana Duarte,
  • Orlando Belo

摘要

Advances in high-throughput technologies have accelerated omics data research. The recent explosion of omics, namely transcriptomic data, has opened new opportunities for the discovery of novel biomarkers with potential to be incorporated into clinical practice. However, due to their extreme complexity, gaining useful insights is particularly challenging. Hence, the application of machine learning techniques on transcriptomic data emerges as a highly promising area for the discovery of new biomarkers. For exploring the potential of these techniques, this paper proposes a novel approach to process gene expression data with the aim of finding candidate gene signatures. Our methodology consists of an ensemble feature selection strategy based on the Boruta, SVM-RFE and LASSO methods, complemented by a second feature selection based on the gene importance calculated by the Random Forest, XGBoost, Support Vector Machine, Logistic Regression and AdaBoost methods. Performing simulations with a dataset of atopic dermatitis patients, our proposal resulted in an 8-gene signature with high AUC (0.839) and accuracy (0.8462) values.