Data deduplication is essential for records in business systems, as duplicate entries can distort analyses, harm relationships, and generate unnecessary costs. As organizations increasingly rely on diverse and large-scale data sources, the challenge of identifying and consolidating duplicate records has grown, making efficient deduplication a critical factor for maintaining data quality and operational efficiency. This paper proposes a hierarchical data deduplication approach to address this industry-wide challenge. The proposed approach, referred to here as the engine, integrates specialized rules, similarity measures, and large language models (LLMs) in sequential layers to analyze data and identify duplications accurately and efficiently. By leveraging a structured combination of techniques, the approach overcomes the limitations of traditional methods, which often depend on isolated and less adaptable solutions. The engine balances performance and accuracy, handling both straightforward and complex cases of duplication. Experimental results using a customer-simulated database demonstrate that integrating these advanced techniques improves adaptability and precision compared to conventional methods. This strategy shows significant potential for large-scale data deduplication, offering a more robust solution to an industry-wide challenge.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a Strategy for Duplication Search Based on Multiple Hierarchically Organized Approaches

  • Cephas A. S. Barreto,
  • Arnaldo S. B. Junior,
  • Leonardo D. G. Melo,
  • Maycon D. R. Santos,
  • Gabriele R. Carvalho,
  • Itamir M. B. Filho,
  • Jose G. R. Neto,
  • Ramon S. Malaquias,
  • Andre M. Gurgel,
  • Jean M. M. Lima

摘要

Data deduplication is essential for records in business systems, as duplicate entries can distort analyses, harm relationships, and generate unnecessary costs. As organizations increasingly rely on diverse and large-scale data sources, the challenge of identifying and consolidating duplicate records has grown, making efficient deduplication a critical factor for maintaining data quality and operational efficiency. This paper proposes a hierarchical data deduplication approach to address this industry-wide challenge. The proposed approach, referred to here as the engine, integrates specialized rules, similarity measures, and large language models (LLMs) in sequential layers to analyze data and identify duplications accurately and efficiently. By leveraging a structured combination of techniques, the approach overcomes the limitations of traditional methods, which often depend on isolated and less adaptable solutions. The engine balances performance and accuracy, handling both straightforward and complex cases of duplication. Experimental results using a customer-simulated database demonstrate that integrating these advanced techniques improves adaptability and precision compared to conventional methods. This strategy shows significant potential for large-scale data deduplication, offering a more robust solution to an industry-wide challenge.