General Methodology for Detecting Fuzzy Duplicates in Electronic Texts with Integrated Mechanisms for Data Confidentiality Preservation
摘要
This article proposes an innovative and improved approach to detecting fuzzy duplicates in electronic texts by combining speed, accuracy, and data confidentiality preservation. It presents a general methodology that meets these requirements and utilizes a combination of shingling and N-gram methods. Its significant advantages are achieved by combining the positive properties of both methods and a specialized algorithm for detecting fuzzy duplicates. The use of shingling allows breaking the text into small fragments and ensuring higher processing speed, while the use of N-grams provides a more accurate analysis of the similarity of text fragments. A key feature of this methodology is the integrated mechanisms for data confidentiality preservation. Measures to protect textual data during their processing and analysis rely on cryptographic encryption and decryption algorithms. Experimental results confirm the high efficiency of the proposed methodology compared to existing approaches. It enables achieving a high accuracy in duplicate detection even for texts with significant variations while maintaining data confidentiality at a high level. The methodology has the potential to become a valuable tool for addressing the issue of growing information volumes and ensuring quality and secure analysis of textual data. It can be applied in various domains, including search engines, plagiarism detection, and applications where safeguarding confidential information is crucial.