An effective imputation approach for handling missing data using intuitionistic fuzzy clustering algorithms
摘要
It is imperative to handle missing data attentively in the preprocessing stage as it may affects the integrity and quality of real-world datasets. However, existing soft clustering-based imputation neglect the underlying non-spherical separability of the data in feature space. This study proposes two robust missing data imputation (MDI) algorithms: Linear Interpolation-based Iterative Intuitionistic Fuzzy C-Means with Euclidean distance (LI-IIFCM) and its weighted variant LI-IIFCM-σ. LI-IIFCM and LI-IIFCM-σ uses linear interpolation for initial imputation followed by iterative IFCM and IFCM-σ, respectively. The approach leverages the soft Davies–Bouldin index to determine the optimal number of clusters and then iteratively refines imputations by minimizing average variation. Experimental analysis and statistical analysis (Friedman Test) on four UCI datasets, using two performance metrics, Mean Absolute Error (MAE) and Root Mean Square Error (RMSE), demonstrate that the proposed algorithms consistently outperform eight existing fuzzy clustering-based MDI algorithms.