错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Genome-wide pervasiveness and localized variation of \(k\)-mer-based genomic signatures in eukaryotes

  • Niousha Sadjadi,
  • Camila P. E. de Souza,
  • Gurjit S. Randhawa,
  • Kathleen A. Hill,
  • Lila Kari

摘要

Genomic signatures–taxon-specific patterns in nucleotide composition–are widely used for taxonomic assignment and comparative genomics, yet their genome-wide pervasiveness across Telomere-to-Telomere assemblies, particularly within functionally diverse and highly repetitive regions, remains undercharacterized. We address this gap with an alignment-free, \(k\) -mer-based analysis using Frequency Chaos Game Representations (FCGRs) across the human genome and three additional eukaryotes from distinct kingdoms. First, by combining qualitative inspection of FCGR landscapes with quantitative distance benchmarking, we show that each species exhibits a stable genomic signature across most chromosomes, with localized departures concentrated in regions enriched for short and long tandem repeats. Then, we introduce two computational pipelines that automatically select a short, contiguous representative genomic segment (500 Kbp) per genome and use it as a proxy to quantify intragenomic variation. Using DSSIM on a [0,1] scale, 80% of 500 Kbp segments in the human genome lie within 0.24 of the representative; segments exceeding this threshold align with tandem-repeat-dense loci. Leveraging these representatives in downstream tasks yields practical gains–for example, one-nearest-neighbor taxonomic classification improves by 7% relative to choosing a random segment. Finally, we provide kCGR-Diff, a graphical tool that enables side-by-side visualization and quantitative comparison of FCGR-based genomic signatures for sample or user-provided sequences, facilitating exploratory analyses of intragenomic variation within and across species. Collectively, our results provide extensive qualitative and quantitative evidence that \(k\) -mer-based genomic signatures are pervasive at genome scale while varying predictably in repeat-dense regions, and they introduce practical methods and software for proxy selection and comparative analysis.