Genomic solutions to metadata challenges: a case study on the 1000 Bull Genomes Project
摘要
Advancements in sequencing technologies have led to an unprecedented availability of whole-genome sequencing data in all life sciences, including livestock research. However, this raises concerns regarding the accuracy of the associated metadata, particularly information on an individual’s subspecies or breed.
ResultsIn this analysis, the 1000 Bull Genomes project was used as an example for a large-scale dataset with structured metadata. We applied a framework combining a query of the NCBI BioSamples database with principal component analysis, admixture analysis, and distance metrics based on the genomic information to assess metadata integrity. The main decrease in metadata quality results from missing breed assignments for 6% (n= 360) of the samples. Moreover, we identified 3% (n= 183) subspecies misassignments and 26 % (n= 1635) of the individuals are involved in the correction of breed assignments. In total, 603 animals receive metadata corrections of the breed or subspecies.
ConclusionOur study highlights the necessity of rigorous metadata assessments, even when utilizing well-established datasets. Furthermore, we emphasize the importance of providing comprehensive and accurate metadata when depositing genomic information in public databases to ensure the integrity and utility of such datasets.