An In-Depth Assessment of Sequence Clustering Software in Bioinformatics
摘要
Sequence clustering software is essential in bioinformatics, yet selecting the most suitable one poses a challenge due to its diverse algorithm design and targeted bioinformatics applications. This paper comprehensively reviewed the developments of most representative sequence clustering software and evaluated 8 representative software based on criteria such as precision, speed, scalability, and memory consumption. This paper divides the clustering software into four aspects: NMI scores greater than 0.95, running time less than 1 min/h, 64 core acceleration exceeding 30 times, and memory consumption less than 3 times the dataset, and summarizes them into a table for user querying. Finally, taking OTU, tree of life building, and metagenomic analysis, as examples, this paper demonstrates how to analyze the requirements of scenarios for clustering software and provides recommendations for selecting the most suitable one based on evaluation results.