Background <p>Advancements in sequencing technologies have led to an unprecedented availability of whole-genome sequencing data in all life sciences, including livestock research. However, this raises concerns regarding the accuracy of the associated metadata, particularly information on an individual’s subspecies or breed.</p> Results <p>In this analysis, the 1000 Bull Genomes project was used as an example for a large-scale dataset with structured metadata. We applied a framework combining a query of the NCBI BioSamples database with principal component analysis, admixture analysis, and distance metrics based on the genomic information to assess metadata integrity. The main decrease in metadata quality results from missing breed assignments for 6% (n= 360) of the samples. Moreover, we identified 3% (n= 183) subspecies misassignments and 26 % (n= 1635) of the individuals are involved in the correction of breed assignments. In total, 603 animals receive metadata corrections of the breed or subspecies.</p> Conclusion <p>Our study highlights the necessity of rigorous metadata assessments, even when utilizing well-established datasets. Furthermore, we emphasize the importance of providing comprehensive and accurate metadata when depositing genomic information in public databases to ensure the integrity and utility of such datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Genomic solutions to metadata challenges: a case study on the 1000 Bull Genomes Project

  • Johanna-Sophie Schlüter-Bartram,
  • Felix Heinrich,
  • Mehmet Gültas,
  • Armin O. Schmitt

摘要

Background

Advancements in sequencing technologies have led to an unprecedented availability of whole-genome sequencing data in all life sciences, including livestock research. However, this raises concerns regarding the accuracy of the associated metadata, particularly information on an individual’s subspecies or breed.

Results

In this analysis, the 1000 Bull Genomes project was used as an example for a large-scale dataset with structured metadata. We applied a framework combining a query of the NCBI BioSamples database with principal component analysis, admixture analysis, and distance metrics based on the genomic information to assess metadata integrity. The main decrease in metadata quality results from missing breed assignments for 6% (n= 360) of the samples. Moreover, we identified 3% (n= 183) subspecies misassignments and 26 % (n= 1635) of the individuals are involved in the correction of breed assignments. In total, 603 animals receive metadata corrections of the breed or subspecies.

Conclusion

Our study highlights the necessity of rigorous metadata assessments, even when utilizing well-established datasets. Furthermore, we emphasize the importance of providing comprehensive and accurate metadata when depositing genomic information in public databases to ensure the integrity and utility of such datasets.