错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Extracting Common DNA Segments from the Complete Genomes of 7538 Viruses and Five Selected Mammals

  • Jing-Doo Wang,
  • Yi-Chun Wang

摘要

This paper aims to extract common DNA consecutive sequences appearing in both of the complete genomes of viruses and that of mammals. With these common DNA sequences as genomic fossils, biologists or virologists can trace possible pathway of genomic evolution across species. To meet the requirement of huge computation to extract common DNA sequences from complete genomes of all viruses and mammals selected, this study adopts one previously developed approach that was based on MapReduce programming model. This study has experiments for extracting common DNA sequences run on a Hadoop cluster containing ten computing nodes. Experimental resources includes the whole genomic sequences of 7,538 viruses and five selected mammals, including \(\textit{Homo sapiens}\) (Human), \(\textit{Pan troglodytes}\) (Chimpanzee), \(\textit{Mus musculus}\) (House Mouse), \(\textit{Rattus norvegicus}\) (Brown Rat) and \(\textit{Sus scrofa}\) (Wild Boar). There are a huge amount of common DNA consecutive sequences extracted and, for simplicity, there are only 26 ones whose lengths are longer than 50 base-pair (bp) selected for illustration. Among above 26, there are 13 reverified as no repetitive sequences that could be seemed as the clues to reinspect the relationship of viruses and mammals. Via cloud computing that can provide with more computing nodes then ten used in this study, it is believed that this approach can handle with more complete genomes of species, and then provide more common genomic fossils to biologists or virologist to verify the potential connections among diverse species in the future.