Knowledge Graph-Based Evaluation Metric for Conversational AI Systems: A Step Towards Quantifying Semantic Textual Similarity
摘要
Machines face difficulty in comprehending the significance of textual information, unlike humans who can easily understand it. The process of semantic analysis aids machines in deciphering the intended meaning of textual information and extracting pertinent data. This, in turn, not only provides valuable insights but also minimizes the need for manual labor. Semantic Textual Similarity in Natural Language processing is one of the most challenging tasks that the research community faces. There have been many traditional methods to evaluate the similarity of two sentences that fall into categories of word-based metrics or embedding based metrics. In natural language understanding (NLU), semantic similarity refers to the degree of likeness or similarity in meaning between two or more pieces of text, words, phrases, or concepts. Evaluating semantic similarity enables the extraction of valuable information, thereby contributing significant insights while minimizing the need for manual labor. The ability of machines to understand the context just like humans do has been a challenging problem for so long. The objective of this research is to introduce a novel evaluation metric for measuring the textual similarity between two texts. The proposed metric will give us NERC scores based on a knowledge-graph approach that can be applied to assess the similarity between texts. The proposed evaluation metrics considers four different features that quantifies the similarity on the level of Nodes, Entities, Relationship and Context. The aggregate score considering the appropriate weights is calculated while considering each of these features in our proposed metrics to generate a final score ranging with least value 0 and maximum value 1. The evaluation was done on the Microsoft MSR Paraphrase Corpus dataset. Along with calculating a NERC score, other scores using already available metrics have also been calculated and reported and the results found during the experiment are comparable to the existing metrics.