Data Structure Identification of Dataset with Missing Values: A Perspective from Granular Computing
摘要
Missing data is frequently encountered in reality, which inevitably poses great challenges to data mining techniques devoted to data structure identification. In view of the fact that missing data (values) generally exhibits high uncertainty, this paper first introduces the concept of information granules, and performs granular imputation on missing data in a more abstract and inclusive way. With the tolerant nature of the information granule to uncertainty, the error of data imputation can be alleviated and the adverse effects on the subsequent research caused by the low-quality data can be reduced. Second, the initial data structure (including granular cluster centers and numerical partition matrix) of the data set with missing values is identified by performing fuzzy clustering on the mixed data set (including numerical and information granules) formed by imputation. Third, the boundaries of granular cluster centers are further optimized by refining the principle of justifiable granularity for mixed data, and a more robust and reliable granular partition matrix is formed subsequently. Finally, by constructing a reconstruction error criterion for mixed data, the critical parameters used in the data clustering process (e.g., the clustering number, the degree of emphasis on the specificity of information granules) are optimized to achieve better data mining results. This paper conducts comprehensive experimental studies on both synthetic and publicly available data sets to show the feasibility and effectiveness of the proposed data structure exploration method.