Fuzzy Clustering Implementations for Big Data in R
摘要
The exponentialDi Perna, V.Ferraro, M.B. growth in data volume, speed, and variety presents both unprecedented opportunities and challenges across diverse domains. Hence, it is imperative to refine the methodologies and to address the intricacies inherent in the analysis of massive datasets. While implementations of the fuzzy k-means algorithm and its variants are provided by numerous R packages, their computation for extensive datasets demand a considerable amount of time. This inefficiency is not unique to such algorithms, as numerous statistical techniques in R lack the ability to leverage modern resources for computational time reduction. The proposed implementations are designed to enhance the efficiency of the fuzzy k-means type clustering algorithms within the R environment through the integration of parallel computing techniques.