Image Data Clustering Based on the Distribution Function
摘要
Image data mining has been widely performed. Grouping images is a challenging task because images contain a lot of color information for an object. The current data clustering challenge is classifying these characteristics. The similarities between the two images describe these the characteristics using the CDF. This study aims to classify images of unstructured data transformed into other forms. The steps of the transformation process are thoroughly discussed in the methodology section. Generally, the images will be processed into a cumulative distribution function (CDF). The clustering algorithm is the k-means method, which efficiently groups data even though it is limited to numeric data types. In this study, a new approach was used: quantifying image data to process it using the k-means algorithm. A challenge in this study is that two or more different images may have the same CDF, resulting in low clustering accuracy due to relatively high error. Therefore, further image processing must ensure that each image has a different distribution function. Another challenge is finding a statistical test to measure the similarity of two images to determine the homogeneity of data within the same cluster. Based on experimental results, these methodologies can classify images with significant accuracy and effectiveness.