Beyond the linear genome: how reference bias threatens preventive medicine and geroscience
摘要
Genotype-first screening compares patient genomes to a standard reference. The GRCh38 linear assembly contains reference minor alleles (RMAs), posing a critical structural vulnerability for automated clinical variant classification. To determine whether RMAs systematically generate false-positive annotations, we conducted an observational case series of 20 healthy adults undergoing preventive genome sequencing. Automated bioinformatic processing was performed using the GRCh38 linear reference to identify “high-impact” variant annotations generated at loci where GRCh38 differs from the population consensus major allele. Among all 20 participants (100%), linear alignment to GRCh38 systematically misclassified functional major alleles as false-positive “high-impact” annotations at three distinct loci (SLC37A4, CIMIP2A, and GPR33). These artifacts occurred solely because analysis software mathematically defined the healthy wild-type state as a deviation from the rare RMA. Cross-referencing with gnomAD confirmed these as population-dominant benign alleles. Furthermore, this inherent reference bias introduces a theoretical risk of false-negative classifications if RMAs mathematically mask true pathogenic variants. Automated pipelines using linear references systematically misclassify healthy alleles as high-impact functional annotations. While downstream population-frequency filters manage false-positive artifacts, this retrospective patching creates an unsustainable bottleneck for population-scale screening. Transitioning to graph-based pangenome references represents a highly promising approach to directly resolve these diagnostic vulnerabilities at the alignment level, though computational and standardization challenges must first be addressed to ensure the accuracy of large-scale aging research.
Graphical Abstract