Multiobjective Interactive Fuzzy Clustering for Gene Expression Data
摘要
Clustering, an unsupervised method for classifying patterns, seeks to organize data points based on their similarities or differences. Common clustering techniques like K-means and Fuzzy C-means often face challenges with local optima, as they focus on optimizing a single cluster validity index. To address this, Genetic Algorithms (GAs) have been used, but their performance across different datasets remains inconsistent. This chapter introduces a new method called Interactive Multiobjective Clustering (IMOC), which adapts and optimizes objective functions during execution. IMOC involves a human decision-maker (DM) in the evaluation process, starting with an initial set of cluster validity indices. To manage DM fatigue, IMOC selectively presents solutions for evaluation, guided by the Non-dominated Sorting Genetic Algorithm II (NSGA-II) and Visual Analysis for Cluster Tendency Assessment (VAT) plots. Evaluations on real-life microarray gene expression datasets, including ”human fibroblasts serum” and ”yeast cell cycle” data, demonstrate IMOC’s superiority over traditional methods like K-means and FCM. Statistical tests confirm IMOC’s significant performance improvement, positioning it as a robust solution for enhancing clustering outcomes across diverse domains.