Ontologies alignment is a critical step in both ontology learning and reuse. Many fuzzy string-matching algorithms have been developed and some are used to perform ontology alignment. However, their empirical evaluation in content-based ontology alignment and their computational performances have not been reported. This study aims to fill this gap through the implementation, empirical analysis and comparison of existing fuzzy string-matching algorithms. Seven fuzzy string-matching algorithms, namely, Jaccard, Jaro-Winkler, LCS, Levenshtein, Cosine Similarity, N-gram, and Damerau-Levenshtein were reviewed and implemented for content-based ontology alignment. The algorithms were tested with a source ontology and four candidate/target ontologies. The algorithms’ performances were evaluated with various metrics including precision, recall, F-measure, and computational performance. Experimental results show that N-gram had the best precision, F-measure, and accuracy, while Jaro-Winkler, Jaccard and Levenshtein performed the worst. Regarding computational performance, Cosine Similarity was the fastest, and N-gram and Damerau-Levenshtein were the slowest algorithms. Memory consumption was comparable, with Dameraul-Levenshtein and Jaccard requiring slightly more space than other algorithms. These findings provide insights into the strengths and limitations of fuzzy string-matching algorithms in the task of content-based ontology alignment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Analysis of Fuzzy String Matching Algorithms for Content-Based Ontology Alignment

  • Mohammed Suleiman Mohammed Rudwan,
  • Jean Vincent Fonou-Dombeu

摘要

Ontologies alignment is a critical step in both ontology learning and reuse. Many fuzzy string-matching algorithms have been developed and some are used to perform ontology alignment. However, their empirical evaluation in content-based ontology alignment and their computational performances have not been reported. This study aims to fill this gap through the implementation, empirical analysis and comparison of existing fuzzy string-matching algorithms. Seven fuzzy string-matching algorithms, namely, Jaccard, Jaro-Winkler, LCS, Levenshtein, Cosine Similarity, N-gram, and Damerau-Levenshtein were reviewed and implemented for content-based ontology alignment. The algorithms were tested with a source ontology and four candidate/target ontologies. The algorithms’ performances were evaluated with various metrics including precision, recall, F-measure, and computational performance. Experimental results show that N-gram had the best precision, F-measure, and accuracy, while Jaro-Winkler, Jaccard and Levenshtein performed the worst. Regarding computational performance, Cosine Similarity was the fastest, and N-gram and Damerau-Levenshtein were the slowest algorithms. Memory consumption was comparable, with Dameraul-Levenshtein and Jaccard requiring slightly more space than other algorithms. These findings provide insights into the strengths and limitations of fuzzy string-matching algorithms in the task of content-based ontology alignment.