The number of biological sequences in the genomic databases, such as the GenBank, has exponentially increased during the past decade. Sequence retrieval systems are required to quickly and efficiently find sequences related to a query sequence. Several comparison algorithms that generally rely upon local string similarities between the query and the database sequences have been widely utilized and accepted as the basis for biosequence retrieval from DNA sequence databases. This chapter describes a new method for sequence comparison based on k-mer word frequency profiles. In this algorithm, the distribution of the k-mer words found on the two sequences, captured by their frequency profiles, are treated as the signatures of sequences. This representation enables us to compare sequence similarity using Shannon’s entropy-based divergence measures.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Biolinguistics

  • Gautam B. Singh

摘要

The number of biological sequences in the genomic databases, such as the GenBank, has exponentially increased during the past decade. Sequence retrieval systems are required to quickly and efficiently find sequences related to a query sequence. Several comparison algorithms that generally rely upon local string similarities between the query and the database sequences have been widely utilized and accepted as the basis for biosequence retrieval from DNA sequence databases. This chapter describes a new method for sequence comparison based on k-mer word frequency profiles. In this algorithm, the distribution of the k-mer words found on the two sequences, captured by their frequency profiles, are treated as the signatures of sequences. This representation enables us to compare sequence similarity using Shannon’s entropy-based divergence measures.