Background <p>Methods for computational discovery of transposable elements (TEs) in DNA sequences are under continual development. Although many TE annotation pipelines have been published, differences in their methods are often opaque and can lead to inconsistent results. As new methods emerge, there is a growing need for an informative and reproducible strategy to evaluate pipeline performance.</p> Results <p>We developed TE_Bench, a user-friendly TE annotation benchmarking workflow that streamlines data generation and visualization for systematic comparison of annotation pipeline performance. TE_Bench can automate simulation of DNA sequences containing artificially evolved TEs from a database or accept user-provided real data to quantify how well test pipelines detect TEs relative to a reference annotation. To accommodate users with varying starting points, TE_Bench is housed as a Snakemake workflow with several options. The data it generates can be used to determine which TE annotation pipeline to use for a specific task, or to inform future improvements to pipelines by revealing shortcomings. We demonstrate the utility of TE_Bench in both contexts by benchmarking EDTA, RepeatModeler2, and Earl Grey using simulated and real DNA sequences. With simulated data, we assess the impact of variables including TE class and nested structure on annotation quality, and with real data, we consider whether read type and mapping method influence downstream annotation. In their default configurations, RepeatModeler2 and Earl Grey perform similarly, outperforming EDTA.</p> Conclusions <p>TE_Bench is an extensible, open-source workflow that supports community-driven TE annotation benchmarking using simulated ground-truth or real genomes. TE_Bench can be accessed at <a href="https://gitub.com/hkania/TE_Bench">https://gitub.com/hkania/TE_Bench</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TE_Bench: a foundational benchmarking workflow for transposable element annotation pipelines

  • Hannah P. Kania,
  • Sierra A. Seifert,
  • Anne D. Yoder

摘要

Background

Methods for computational discovery of transposable elements (TEs) in DNA sequences are under continual development. Although many TE annotation pipelines have been published, differences in their methods are often opaque and can lead to inconsistent results. As new methods emerge, there is a growing need for an informative and reproducible strategy to evaluate pipeline performance.

Results

We developed TE_Bench, a user-friendly TE annotation benchmarking workflow that streamlines data generation and visualization for systematic comparison of annotation pipeline performance. TE_Bench can automate simulation of DNA sequences containing artificially evolved TEs from a database or accept user-provided real data to quantify how well test pipelines detect TEs relative to a reference annotation. To accommodate users with varying starting points, TE_Bench is housed as a Snakemake workflow with several options. The data it generates can be used to determine which TE annotation pipeline to use for a specific task, or to inform future improvements to pipelines by revealing shortcomings. We demonstrate the utility of TE_Bench in both contexts by benchmarking EDTA, RepeatModeler2, and Earl Grey using simulated and real DNA sequences. With simulated data, we assess the impact of variables including TE class and nested structure on annotation quality, and with real data, we consider whether read type and mapping method influence downstream annotation. In their default configurations, RepeatModeler2 and Earl Grey perform similarly, outperforming EDTA.

Conclusions

TE_Bench is an extensible, open-source workflow that supports community-driven TE annotation benchmarking using simulated ground-truth or real genomes. TE_Bench can be accessed at https://gitub.com/hkania/TE_Bench.