<p>As next-generation sequencing technologies produce deeper genome coverages at lower costs, there is a critical need for reliable computational host DNA removal in metagenomic data. We find that insufficient host filtration using prior human genome references can introduce false sex biases and inadvertently permit flow-through of host-specific DNA during bioinformatic analyses, which could be exploited for individual identification. To address these issues, we introduce and benchmark three host filtration methods of varying throughput, with concomitant applications across low biomass samples such as skin and high microbial biomass datasets including fecal samples. We find that these methods are important for obtaining accurate results in low biomass samples (e.g., tissue, skin). Overall, we demonstrate that rigorous host filtration is a key component of privacy-minded analyses of patient microbiomes and provide computationally efficient pipelines for accomplishing this task on large-scale datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Incomplete human reference genomes can drive false sex biases and expose patient-identifying information in metagenomic data

  • Caitlin Guccione,
  • Lucas Patel,
  • Yoshihiko Tomofuji,
  • Daniel McDonald,
  • Antonio Gonzalez,
  • Gregory D. Sepich-Poore,
  • Kyuto Sonehara,
  • Mohsen Zakeri,
  • Yang Chen,
  • Amanda Hazel Dilmore,
  • Neil Damle,
  • Sergio E. Baranzini,
  • George Hightower,
  • Teruaki Nakatsuji,
  • Richard L. Gallo,
  • Ben Langmead,
  • Yukinori Okada,
  • Kit Curtius,
  • Rob Knight

摘要

As next-generation sequencing technologies produce deeper genome coverages at lower costs, there is a critical need for reliable computational host DNA removal in metagenomic data. We find that insufficient host filtration using prior human genome references can introduce false sex biases and inadvertently permit flow-through of host-specific DNA during bioinformatic analyses, which could be exploited for individual identification. To address these issues, we introduce and benchmark three host filtration methods of varying throughput, with concomitant applications across low biomass samples such as skin and high microbial biomass datasets including fecal samples. We find that these methods are important for obtaining accurate results in low biomass samples (e.g., tissue, skin). Overall, we demonstrate that rigorous host filtration is a key component of privacy-minded analyses of patient microbiomes and provide computationally efficient pipelines for accomplishing this task on large-scale datasets.