Biological Sequence Clustering: Novel Approaches and a Comparative Study
摘要
The application of clustering tools plays an essential role in the examination of biological sequences. While most available tools utilize greedy and hierarchical algorithms, spectral clustering has recently emerged as a game changer in this domain. In prior research, a clustering approach based on spectral clustering was introduced. Although it showed superiority over cutting-edge clustering methods for divergent sequences, it had limitations such as high computational demands, inability to manage extensive datasets including sequences of different lengths, and incompatibility with some dataset types. Most of these issues were addressed in this study. Initially, a significant enhancement in the speed of pairwise affinity computations was realized. Subsequently, several innovative clustering methods, such as MOTIFS-based clustering, which are utilized for clustering different data types, were implemented and adapted for biological sequence clustering. Furthermore, a new clustering technique, CHAINS, was introduced. Ultimately, a detailed qualitative evaluation of the examined clustering methods on biological sequences was performed. The CHAINS method outperformed its predecessors in terms of speed and clustering performance, particularly with datasets that contain large genomes. The outcomes from the experiments were compared to those from leading clustering tools, which the CHAINS method also outperformed in clustering hybrid and highly divergent datasets. The source code for the clustering tool, which incorporates all advancements and examined methodologies, is freely available online.