Microbial community analysis based on high-throughput sequencing data poses significant statistical challenges. A wide array of methods exists to quantify between-sample diversity, yet the biological interpretation of such metrics remains unclear or inconsistent. In particular, the “double-zero problem”, the shared absence of a taxon in two samples, introduces semantic ambiguity that may lead to misleading conclusions. To evaluate this issue, we applied 25 combinations of data transformations and distance formulas to 16 microbial abundance tables from public 16S rRNA datasets. Two clusters of highly correlated distances (r > 0.85) were identified, with 16 and 4 metrics respectively. Distance matrices from a single project were selected as a case study, and multiple statistical tools were used to explore the structural impact of double zeros. Manhattan, Euclidean, and Binomial distances showed strong association with double-zero proportions, and are therefore not recommended for reporting beta diversity. However, they may serve as useful indicators of double-zero prevalence. Finally, we propose a conceptual model for representing distance matrices, which may aid in the development of more robust analytical strategies in future beta diversity studies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empirical Analysis of the Double-Zero Problem in Distances Between Genomic Samples in Microbial Abundance Matrices: A Case Study

  • Sofia Oiene,
  • Nicolás Gamboa,
  • Rosario Taussig,
  • Debora Chan,
  • Ignacio Cassol

摘要

Microbial community analysis based on high-throughput sequencing data poses significant statistical challenges. A wide array of methods exists to quantify between-sample diversity, yet the biological interpretation of such metrics remains unclear or inconsistent. In particular, the “double-zero problem”, the shared absence of a taxon in two samples, introduces semantic ambiguity that may lead to misleading conclusions. To evaluate this issue, we applied 25 combinations of data transformations and distance formulas to 16 microbial abundance tables from public 16S rRNA datasets. Two clusters of highly correlated distances (r > 0.85) were identified, with 16 and 4 metrics respectively. Distance matrices from a single project were selected as a case study, and multiple statistical tools were used to explore the structural impact of double zeros. Manhattan, Euclidean, and Binomial distances showed strong association with double-zero proportions, and are therefore not recommended for reporting beta diversity. However, they may serve as useful indicators of double-zero prevalence. Finally, we propose a conceptual model for representing distance matrices, which may aid in the development of more robust analytical strategies in future beta diversity studies.