错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph Similarity Join (GSJ) Approach to Detect Near Duplicate Text Documents

  • Prathi Naveena,
  • Sandeep Kumar Dash

摘要

With the profligate progress of graph-based data in numerous areas essential to observe the similarities among the graph pairs is indispensable to speed up the search outcomes. The existence of near-duplicate text documents plays an imperative role in performance degradation while integrating information from diverse sources. The main issue is determining the similarity among the graphs from massive datasets. The principal goal of this paper is to enhance search optimization with improved indexing and condensed pairwise comparisons with the help of the proposed Graph Similarity Self Join of Documents (GSSJD) technique is formed by uniting the prefix filtering strategy to the traditional inverted index method. Experiments were carried out on the dataset to authenticate that our proposed approach illustrates the optimized results in speed and time.