Data Analysis of Discrete-Valued Models for Genetic Sequences
摘要
Two families of parsimonious Markov models for statistical analysis of genetic sequences are considered. The first family of conditionally nonlinear autoregressive (CNAR) models is useful in such problems as detection and description of deep Markov dependencies in long genetic sequences and recognition of protein coding regions. The second family of maximum entropy models (MEM) is useful in problem of discrimination of special human DNA signals (donor and acceptor splice sites, start codon, stop codon) from decoys.