An Efficient Methodology for Identifying the Similarity Between Languages with Levenshtein Distance
摘要
This article compares the multilingual texts that are used for bilingual lexicon extraction and plagiarism detection. A collection of related sentences and sentences that are translations of one another, a parallel corpus, is used to compare multilingual content. In order to determine the similarity between the sentences and words in a multilingual content, Levenshtein Distance based String Similarity technique is used along with three other techniques such as Sequence matcher, Fuzzy-Wuzzy (Ratio), and Spacy Similarity techniques in the literature. The Comparative study is presented in this article to compare similar kind of work implemented with proposed techniques. The Levenshtein Distance based string similarity technique out performs in terms of Accuracy compared to Sequence matcher, Fuzzy-Wuzzy and Spacy similarity techniques for the identification of similarity between the languages.