Comparative Study of Different Document Similarity Measures and Models
摘要
Document similarity refers to an approach of measuring how two or more documents look alike in terms of their content or structure. Document similarity algorithms are used to determine the degree of resemblance or relatedness between various documents. Document similarity plays a pivotal role in a wide range of tasks involving natural language processing, information retrieval, recommender systems and duplicates detection. In this paper, we will be studying and compare the similarity score of documents using different document similarity measures and models like cosine similarity, Euclidean distance, Jaccard similarity, Latent Semantic Analysis (LSA), Latent Dirichlet Allocation (LDA), Bidirectional Encoder Representations from Transformers (BERTs), etc.