Processing-Bias Correction with DEBIAS-M Improves Cross-Study Generalization of Microbiome-Based Prediction Models
摘要
Microbiome profiling exhibits strong study- and batch-specific effects, impeding the identification of signals that are reproducible across studies and the development of generalizable prediction models. Prior work has attributed this to biases introduced during experimental protocols [1], with factors such as the type of DNA extraction kit affecting the efficiency of extracting and sequencing different microbes [2, 3]. While existing batch-correction methods show benefit in microbiome analysis [4–7], many make strong parametric assumptions, which do not necessarily apply in this data, or require the use of the outcome variable, which risks overfitting [8]. Lastly and importantly, the transformations performed to the data are largely non-interpretable, e.g., introducing counts to features that were initially very sparse. Here, we present DEBIAS-M (Domain adaptation with phenotype Estimation and Batch Integration Across Studies of the Microbiome), an interpretable framework for processing-bias inference, batch correction, and domain adaptation in microbiome studies. DEBIAS-M learns bias-correction factors for each microbe in each batch that simultaneously minimize batch effects and maximize cross-study associations with phenotypes. Using benchmarks, including HIV classification from gut microbiome data, we demonstrate that DEBIAS-M outperforms alternative batch-correction methods commonly used in the field. Overall, we show that DEBIAS-M facilitates better modeling of microbiome data and identification of signals that are reproducible across studies.