Phylogeny Reconstruction Using \(k-mer\) Derived Transition Features
摘要
Traditional, sequence alignment-based (AB) methods are time-consuming and show \(NP-hard\) complexity for large sequences. That is why alignment-free (AF) approaches become popular among scientists due to low time complexity. Besides, sequence analysis is crucial for gene finding, modification, new variety development, etc. However, existing AF algorithms utilized \(k-mer\) count, histogram, chaos game representation (CGR), etc., but these have low accuracy rates. Therefore, in this research, a novel \(k-mer\) derived transition spatial feature along with standard deviation and the median is used for representing a sequence. Median and standard deviation are extracted from first-order derivatives of \(k-mer\) positions, and the transition is derived from second-order derivatives. The method is tested in six challenging benchmark datasets from different viewpoints and achieved top-rank accuracy for all datasets and saves 23–99% memory consumption. Here, the top accuracy indicates that extracted features are highly efficient to represent the inherent property of a sequence. Moreover, the time consumption is close to state-of-the-art methods and a thousand times faster than the existing MEGA tool. Therefore, industries can use this method with conviction.