Gene expression data mining by hybrid biclustering with improved GA and BA
摘要
Accurately clustering gene expression data is essential for understanding biological processes and disease mechanisms. However, this remains challenging due to the datasets’ inherent high dimensionality, complexity, and noise. To overcome the limitations of conventional clustering approaches, this study presents a dual clustering method that integrates enhanced heuristic algorithms. The study improves the shortcomings of the genetic algorithm and the bat algorithm in the optimization process. The two are then merged to form a dual clustering method for gene expression data. The outcomes revealed that both this improved genetic algorithm and the improved bat algorithm showed higher convergence speed and optimal solution solving accuracy than the other heuristic algorithms in the computation of single-peak function and multi-peak function calculations. In the dual clustering visualization results, the three inter-cluster distances of the dual clustering results of the research method were farther and the intra-cluster distances were closer. The geometric mean of hybrid dual clustering was 0.99, silhouette coefficient value was 1.0, Davies-Bouldin index was 0.2, and adjusted rand index mean was 0.92, all of which were better than other dual clustering methods. The results indicated that the proposed dual clustering method achieved the clustering effect of high inter-cluster variability and high intra-cluster similarity. The proposed dual clustering method improves the efficiency and accuracy of gene expression data analysis and provides a powerful technical support for bioinformatics research.