Background <p>An accurate genome annotation is essential in many contexts, including RNA sequencing studies. Annotations include known genes and isoforms, detailing their location (chromosome, start, and end) and coding sequence, among other important metadata.</p> Results <p>We characterized changes in human Ensembl annotations from 2014 to 2023 and the important gains in our biological understanding in recent years. While generally gene and isoform annotations increased (2014: 58,812 genes ; 2023: 62,710), some years dropped (e.g., 2016). A similar pattern exists for the gene and isoform biotypes; both 2015 (19,825) and 2017 (19,828) have fewer genes annotated as protein-coding than 2014 (19,953) and 2016 (19,961)— 2023 has the most (20,048). <i>PCBP1-AS1</i> had the most annotated isoforms (296). We quantified expression for isoforms that were new between 2019 and 2023 across nine GTEx tissues (58 samples) to demonstrate our significant gains in understanding recently. We saw 2,054 of these ‘new’ isoforms expressed in cerebellar hemisphere (594 in liver). For many genes, we saw that the relative expression of the ‘new’ isoforms was much greater than the previously known isoforms.</p> Conclusions <p>This study demonstrates the importance of an accurate genome annotation to truly understand the underlying complexity of biology that is often oversimplified by ignoring transcriptional complexity.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Genome annotations matter: characterizing Ensembl hg38 annotations from 2014 to 2023

  • Madeline L. Page,
  • Mark E. Wadsworth,
  • Bernardo Aguzzoli Heberle,
  • David W. Fardo,
  • Mark T. W. Ebbert

摘要

Background

An accurate genome annotation is essential in many contexts, including RNA sequencing studies. Annotations include known genes and isoforms, detailing their location (chromosome, start, and end) and coding sequence, among other important metadata.

Results

We characterized changes in human Ensembl annotations from 2014 to 2023 and the important gains in our biological understanding in recent years. While generally gene and isoform annotations increased (2014: 58,812 genes ; 2023: 62,710), some years dropped (e.g., 2016). A similar pattern exists for the gene and isoform biotypes; both 2015 (19,825) and 2017 (19,828) have fewer genes annotated as protein-coding than 2014 (19,953) and 2016 (19,961)— 2023 has the most (20,048). PCBP1-AS1 had the most annotated isoforms (296). We quantified expression for isoforms that were new between 2019 and 2023 across nine GTEx tissues (58 samples) to demonstrate our significant gains in understanding recently. We saw 2,054 of these ‘new’ isoforms expressed in cerebellar hemisphere (594 in liver). For many genes, we saw that the relative expression of the ‘new’ isoforms was much greater than the previously known isoforms.

Conclusions

This study demonstrates the importance of an accurate genome annotation to truly understand the underlying complexity of biology that is often oversimplified by ignoring transcriptional complexity.