With the development of the big data era, the volume of data is on an explosive growth trend. It will be quite a challenge to mine useful and valuable information from tons of data. It is simply impossible to accomplish such mining on a single computer. Hence, it becomes essential to utilize computer clusters for data mining. Infinitely scalable storage and computational power will solve the data mining problem posed by the data explosion. Therefore, it becomes important to identify the shortcomings of existing parallelization algorithms and improve them, and to propose an algorithm that can improve computational efficiency, reduce storage burden, and achieve load balancing. At the same time, we can utilize big data technology: MapReduce and HDFS. Using HDFS for file sharding to ensure load balancing and MapReduce principle for data processing to reduce the interaction between servers. The efficiency of the improved algorithm can be verified by conducting experiments of frequent itemset mining on a cluster of computers and recording the runtime. The scalability of the algorithm in a cluster environment can be demonstrated by performing frequent itemset mining on different cluster sizes and recording the running time.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Big Data Analytics: Advanced Parallelization Techniques and Adaptive Resource Allocation Frameworks

  • Baokui Liao,
  • S. B. Goyal,
  • Anand Singh Rajawat,
  • A. Z. M. Ibrahim

摘要

With the development of the big data era, the volume of data is on an explosive growth trend. It will be quite a challenge to mine useful and valuable information from tons of data. It is simply impossible to accomplish such mining on a single computer. Hence, it becomes essential to utilize computer clusters for data mining. Infinitely scalable storage and computational power will solve the data mining problem posed by the data explosion. Therefore, it becomes important to identify the shortcomings of existing parallelization algorithms and improve them, and to propose an algorithm that can improve computational efficiency, reduce storage burden, and achieve load balancing. At the same time, we can utilize big data technology: MapReduce and HDFS. Using HDFS for file sharding to ensure load balancing and MapReduce principle for data processing to reduce the interaction between servers. The efficiency of the improved algorithm can be verified by conducting experiments of frequent itemset mining on a cluster of computers and recording the runtime. The scalability of the algorithm in a cluster environment can be demonstrated by performing frequent itemset mining on different cluster sizes and recording the running time.