The discovery of genes that code for a specific enzymatic activity is important in various fields of life science and provides valuable biotechnological tools. Many genes that contribute to the production of secondary metabolites and specialized metabolic pathways are still not identified. Due to the great diversity of metabolic functions found in nature and their rapid evolutionary adaptation, we need precise but high-throughput approaches for a targeted search based on minimal prior knowledge. In this chapter, we describe a transcriptomics pipeline that was used to search for candidate genes coding for a specific enzymatic activity in a nonmodel species. We generated and combined short- and long-read transcriptomic data to obtain reliable full-length transcript sequences along with information on allelic variation, isoform expression, and condition-specific expression. Based on protein domain annotations of coding sequences and transcriptomic data, we selected candidate genes for activity assays. We provide detailed instructions for analysis and quality control steps in our pipeline that can be applied to other biological questions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Combining Short- and Long-Read Transcriptomes for Targeted Enzyme Discovery

  • Mojca Juteršek,
  • Marko Petek,
  • Špela Baebler

摘要

The discovery of genes that code for a specific enzymatic activity is important in various fields of life science and provides valuable biotechnological tools. Many genes that contribute to the production of secondary metabolites and specialized metabolic pathways are still not identified. Due to the great diversity of metabolic functions found in nature and their rapid evolutionary adaptation, we need precise but high-throughput approaches for a targeted search based on minimal prior knowledge. In this chapter, we describe a transcriptomics pipeline that was used to search for candidate genes coding for a specific enzymatic activity in a nonmodel species. We generated and combined short- and long-read transcriptomic data to obtain reliable full-length transcript sequences along with information on allelic variation, isoform expression, and condition-specific expression. Based on protein domain annotations of coding sequences and transcriptomic data, we selected candidate genes for activity assays. We provide detailed instructions for analysis and quality control steps in our pipeline that can be applied to other biological questions.