Probabilistic Data Structure Using Hashing Technique for Big Data Security De-duplication in Cloud Environment
摘要
The exponential growth of data has seen unexpected elevation with time. Technology such as cloud computing and modern data analysis methods such as Hadoop and Map Reduce are helping to store and handle this large volume of data. However, the storage and security are major concerns. Data redundancy is a common problem in cloud and big data. De-duplication methods are commonly used to improve the efficiency of storage in cloud and can save network bandwidth. These, however, have potential risks such as privacy preservation and other cyber-attacks. In this research paper, a secured scheme for de-duplication is proposed using one of the Probabilistic Data Structure, and the Bloom filter is proposed to achieve more efficiency of the proof of ownership which can realize quick detection of redundant data chunks and blocks and thus bringing down the false positive rate. The proposed method also concentrates on minimizing the cyber-attacks through double encryption of file when uploaded in cloud. The experimental results prove that the proposed method has less computational time and minimized false positive rate when compared to the base schemes.