错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study of Different Document Similarity Measures and Models

  • Anshika Singh,
  • Sharvan Kumar Garg

摘要

Document similarity refers to an approach of measuring how two or more documents look alike in terms of their content or structure. Document similarity algorithms are used to determine the degree of resemblance or relatedness between various documents. Document similarity plays a pivotal role in a wide range of tasks involving natural language processing, information retrieval, recommender systems and duplicates detection. In this paper, we will be studying and compare the similarity score of documents using different document similarity measures and models like cosine similarity, Euclidean distance, Jaccard similarity, Latent Semantic Analysis (LSA), Latent Dirichlet Allocation (LDA), Bidirectional Encoder Representations from Transformers (BERTs), etc.