Query-document vocabulary mismatch represents the gap between a query’s terms and the index terms used for document retrieval. It is a significant challenge that affects severely the performance of search algorithms. Our Ph.D. focuses on building a semantic layer that can be shared by both document index terms as well as query terms in order to overcome this problem. In this paper we focus on expanding queries using aligned keyphrases. We show that state-of-the-art keyphrase generation models do improve retrieval but at the cost of an increased vocabulary mismatch. To reduce this effect, we project, using sentence-transformers, the generated keyphrases to their closest representative term from the indexed vocabulary. However, the original set consists of author-assigned annotations which may suffer from issues such as duplication and misspelling. Through the processing of these annotations, we are able to reduce the search space for query-document alignment. We repeat this experiment on keyphrases extracted by tf-idf and demonstrate significant improvements over the author keyphrases, effectively bridging the vocabulary gap and enhancing search relevance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Sentence-Transformers to Overcome Query-Document Vocabulary Mismatch in Information Retrieval

  • Saber Zahhar,
  • Nédra Mellouli,
  • Christophe Rodrigues

摘要

Query-document vocabulary mismatch represents the gap between a query’s terms and the index terms used for document retrieval. It is a significant challenge that affects severely the performance of search algorithms. Our Ph.D. focuses on building a semantic layer that can be shared by both document index terms as well as query terms in order to overcome this problem. In this paper we focus on expanding queries using aligned keyphrases. We show that state-of-the-art keyphrase generation models do improve retrieval but at the cost of an increased vocabulary mismatch. To reduce this effect, we project, using sentence-transformers, the generated keyphrases to their closest representative term from the indexed vocabulary. However, the original set consists of author-assigned annotations which may suffer from issues such as duplication and misspelling. Through the processing of these annotations, we are able to reduce the search space for query-document alignment. We repeat this experiment on keyphrases extracted by tf-idf and demonstrate significant improvements over the author keyphrases, effectively bridging the vocabulary gap and enhancing search relevance.