An unsupervised tool for biomarker discovery and cancer subtyping applied to glioblastoma
摘要
High-dimensional omics data often contain more variables than observations, which can lead to overfitting and negatively impact the results of classical data analysis methods. To address the issue, supervised variable selection methods are often used, incorporating penalty terms into the model. While effective for selecting task-specific variables, this approach may not preserve the overall dataset structure for multiple downstream analyses. This study aims to evaluate unsupervised variable selection approaches and introduce a novel tool that improves data interpretability while maintaining biological information.
ResultsWe assessed multiple unsupervised variable selection techniques to identify a representative subset of the original dataset. Based on this evaluation, we developed TRIM-IT, a computational tool that integrates unsupervised variable selection, clustering, survival analysis, and differential gene expression analysis. TRIM-IT was applied to glioblastoma transcriptomics data, uncovering three distinct patient clusters. These clusters correlated with tumor histology, exhibited significantly different survival outcomes, and revealed molecular profiles that suggest potential biomarker candidates.
ConclusionTRIM-IT provides a novel approach for analyzing high-dimensional omics data while preserving key biological insights. Its ability to identify meaningful patient subgroups and molecular signatures highlights its applicability across various biomedical research contexts. The tool is implemented in R and the code is publicly available for reproduction and adaptation to other studies.