Applications of Computational and Data Sciences in Metabolomics
摘要
Metabolomics is a field of systems biology that involves the analysis of small molecule metabolites present in tissues or biofluids. Thanks to the development of mass spectrometry (MS) and nuclear magnetic resonance (NMR) spectroscopy, metabolomic data have rapidly accumulated, which requires robust tools to interpret the vast amount of metabolomic data into biological meanings. Current tools for metabolomic analysis like MetaboAnalyst, XCMS, or CAMERA mostly integrate both statistical models like univariate and multivariate analysis, machine learning techniques like support vector machines (SVM), random forests (RF), and even neural networks, and visualization tools. From which, metabolites and metabolic pathways are detected for deeply understanding about cellular metabolism, regulatory networks, and metabolic phenotypes. The roles of data science and computational science are clearly demonstrated in every step of metabolomic data processing, including data collection and organization, preprocessing, quality control, annotation, model building and validation, and statistical analysis as well as data visualization. In this review, our aim is to provide a comprehensive and robust view of the development history and the current status of statistical and bioinformatic tools in data science and computational science applied for metabolomics.