Public Domain Databases: A Gold Mine for Identification and Genome Reconstruction of Plant Viruses and Viroids
摘要
Plant viruses deprive the host due to their replication in the infected cells by hijacking the host machinery and thereby affecting the quality and quantity of economic produce. In recent years, Next Generation Sequencing (NGS) technologies are widely adopted for plant viral detection as it facilitates identification of novel viruses and multiple infections which otherwise could not be detected using traditional methods. On the other hand, reducing costs of plant genome and transcriptome sequencing projects lead to the deposition of huge sequence data in the NCBI Sequence Read Archive (SRA) and the Transcriptome Shotgun Assembly (TSA) databases. In addition to host reads and contigs, these datasets may also contain the inadvertently co-isolated viral sequences. In general, to explore the virome of a targeted plant species, datasets generated from that particular plant and available in SRA database are analysed while the TSA database is explored to identify novel viral sequences of a targeted virus group in several plant species. The workflow for identification and genome reconstruction of plant viruses and viroids using public domain databases is discussed. Though viral/viroid genome sequences recovered solely from metagenomic data could be regarded as bona fide ones, wherever possible experimental studies can be undertaken to validate their presence in biological samples and to complete the genome sequences recovered from public domain datasets.