An Assessment of the Utilization of MapReduce Expectation-Maximization Grouping Using Ensemble and Accumulation of Bootstrap for Big Data Analysis
摘要
The resolution of the clustering problem holds significant importance in the realm of research and information discovery. In the realm of big data, the task of categorizing similar data poses a significant challenge owing to the vast volume of data involved. Numerous clustering techniques were devised in previous studies. To address the aforementioned limitations, a novel technique called Bootstrap Aggregated Map Reduce Expectation-Maximization Clustering (BAMEC) is introduced. The BAMEC approach initially utilizes the comprehensive Brazilian E-Commerce Public Dataset as its primary input. The BAMEC approach utilizes user input to create a certain amount of bootstrap specimens from a large input dataset. Following this, an approach known as BAMEC (Bootstrap Aggregating Map-Reduced Expectation–Maximization Clustering) is employed to develop the optimal amount of Map Reduction based Expectation–Maximization (MEM) clustering algorithms for each bootstrap sample that has been generated. Subsequently, an approach known as BAMEC is employed to amalgamate the outcomes of all MEM clustering algorithms, followed by the application of a voting mechanism. Finally, the BAMEC technique effectively groups similar data by taking into account the outcomes of majority voting. The BAMEC approach employs various measures, including clustering accuracy, computing cost, false alarm rate, and space complexity, to evaluate the experimental procedure. The experimental findings demonstrate that the BAMEC approach has the capability to improve the reliability of clustering and also decrease the computational burden of big data analytics in comparison to existing approaches.