<p>Making sense of whole-genome polymorphism data is challenging, but it is essential for overcoming the biases in SNP data. Here we analyze 27 genomes of <i>Arabidopsis thaliana</i> to illustrate these issues. Genome size variation is mostly due to tandem repeat regions that are difficult to assemble. However, while the rest of the genome varies little in length, it is full of structural variants, mostly due to transposon insertions. Because of this, the pangenome coordinate system grows rapidly with sample size and ultimately becomes 70% larger than the size of any single genome, even for <i>n</i> = 27. Finally, we show how short-read data are biased by read mapping. SNP calling is biased by the choice of reference genome, and both transcriptome and methylome profiling results are affected by mapping reads to a reference genome rather than to the genome of the assayed individual.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A comparison of 27 Arabidopsis thaliana genomes and the path toward an unbiased characterization of genetic polymorphism

  • Anna A. Igolkina,
  • Sebastian Vorbrugg,
  • Fernando A. Rabanal,
  • Hai-Jun Liu,
  • Haim Ashkenazy,
  • Aleksandra E. Kornienko,
  • Joffrey Fitz,
  • Max Collenberg,
  • Christian Kubica,
  • Almudena Mollá Morales,
  • Benjamin Jaegle,
  • Travis Wrightsman,
  • Vitaly Voloshin,
  • Alexander D. Bezlepsky,
  • Victor Llaca,
  • Viktoria Nizhynska,
  • Ilka Reichardt,
  • Ilja Bezrukov,
  • Christa Lanz,
  • Felix Bemm,
  • Pádraic J. Flood,
  • Sileshi Nemomissa,
  • Angela Hancock,
  • Ya-Long Guo,
  • Paul Kersey,
  • Detlef Weigel,
  • Magnus Nordborg

摘要

Making sense of whole-genome polymorphism data is challenging, but it is essential for overcoming the biases in SNP data. Here we analyze 27 genomes of Arabidopsis thaliana to illustrate these issues. Genome size variation is mostly due to tandem repeat regions that are difficult to assemble. However, while the rest of the genome varies little in length, it is full of structural variants, mostly due to transposon insertions. Because of this, the pangenome coordinate system grows rapidly with sample size and ultimately becomes 70% larger than the size of any single genome, even for n = 27. Finally, we show how short-read data are biased by read mapping. SNP calling is biased by the choice of reference genome, and both transcriptome and methylome profiling results are affected by mapping reads to a reference genome rather than to the genome of the assayed individual.