Benchmark Study on Supervised Relevance-Redundancy Assessment for Feature Selection in Genomic Data
摘要
Single variant Genome-Wide Association Studies (GWAS) are the most common data-driven approach to discover genetic variants associated to phenotypes. However, in complex diseases, single variants often have no effect unless they coexist with other variants. Conversely, machine learning (ML) can model potential interactions among variants, potentially addressing the missing heritability problem in these diseases. Nevertheless, the curse of dimensionality must be considered, given the tremendous number of variants in genomic datasets, requiring feature selection techniques to reduce the number of features. This study aims at a preliminary benchmark of the Relevance-Redundancy assessment (ReRa) feature selection method using a public genetic dataset of a Parkinson’s cohort. Obtained results demonstrated that ReRa can achieve performances comparable to common filter-based feature selection techniques, with the benefit of building simpler models with fewer features.