<p>Long-read metatranscriptomics is a powerful and cost-effective technology for elucidating the genetic diversity and expression dynamics of active eukaryotic microorganisms by characterizing full-length transcripts. However, its potential has been limited by the lack of high-quality reference genomes and high sequencing error rates. We present Fungen, a reference-free tool that constructs accurate transcripts from long-read metatranscriptomic data through read clustering and error correction. Fungen achieves superior accuracy in transcript determination while significantly reducing memory usage and offering a 22 to 56-fold speed improvement over existing methods. This novel approach overcomes the challenges posed by sequence similarity among closely related species, enabling the analysis of deeply sequenced metatranscriptomes by generating reliable gene clusters and accurate sequences. Two applications showcase Fungen’s capabilities to perform high-resolution taxonomic assignments and gene profiling in marine direct RNA datasets, as well as resolving reliable annotation identities in full-length rRNA targeted sequencing datasets. When applied to soil metatranscriptomic data, Fungen offers valuable insights into the <i>in situ</i> fungal composition and gene expression dynamics, revealing specialized life strategies of plant-pathogenic fungi in soil environments. Overall, Fungen provides a fast, scalable, and accurate solution for analyzing complex metatranscriptomic datasets, paving the way for a comprehensive understanding of eukaryotic diversity and function from long-read sequencing data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fungen: clustering and correcting long-read metatranscriptomic data for exploring eukaryotic microorganisms

  • Weiwei Zhang,
  • Xiang Jennie Li,
  • Fang Liu,
  • Jie Zhang,
  • Jianqing Tian,
  • Yuan Gao

摘要

Long-read metatranscriptomics is a powerful and cost-effective technology for elucidating the genetic diversity and expression dynamics of active eukaryotic microorganisms by characterizing full-length transcripts. However, its potential has been limited by the lack of high-quality reference genomes and high sequencing error rates. We present Fungen, a reference-free tool that constructs accurate transcripts from long-read metatranscriptomic data through read clustering and error correction. Fungen achieves superior accuracy in transcript determination while significantly reducing memory usage and offering a 22 to 56-fold speed improvement over existing methods. This novel approach overcomes the challenges posed by sequence similarity among closely related species, enabling the analysis of deeply sequenced metatranscriptomes by generating reliable gene clusters and accurate sequences. Two applications showcase Fungen’s capabilities to perform high-resolution taxonomic assignments and gene profiling in marine direct RNA datasets, as well as resolving reliable annotation identities in full-length rRNA targeted sequencing datasets. When applied to soil metatranscriptomic data, Fungen offers valuable insights into the in situ fungal composition and gene expression dynamics, revealing specialized life strategies of plant-pathogenic fungi in soil environments. Overall, Fungen provides a fast, scalable, and accurate solution for analyzing complex metatranscriptomic datasets, paving the way for a comprehensive understanding of eukaryotic diversity and function from long-read sequencing data.