The automated construction of topic taxonomies can benefit numerous applications, including web search, recommendation, and knowledge discovery. One of the major advantages of automatic taxonomy construction is the ability to capture corpus-specific information and adapt to different scenarios. Traditionally, many approaches have relied on word2vec to create vector representations of keywords, but this method comes with notable shortcomings. While word2vec focuses mainly on syntactic relationships, it often falls short in capturing deeper semantic connections. It struggles with words that have multiple meanings and doesn’t thoroughly account for the context in longer pieces of text, which can result in less effective taxonomy structures. In this paper, we propose a novel approach that addresses this issue by utilizing large language models (LLMs) to generate similarity scores between pairs of keywords or key phrases. We then use these similarity scores to train a sentence embedding model that can capture semantic relationships and context-dependent meanings. Unlike word2vec, our approach takes advantage of the detailed contextual embeddings provided by LLMs, leading to more accurate and meaningful taxonomies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Taxonomy Construction via Sentence Embedding Models Trained with LLM-Generated Similarity Scores

  • Chi-Cuong Tran,
  • Tuan-Dung Cao

摘要

The automated construction of topic taxonomies can benefit numerous applications, including web search, recommendation, and knowledge discovery. One of the major advantages of automatic taxonomy construction is the ability to capture corpus-specific information and adapt to different scenarios. Traditionally, many approaches have relied on word2vec to create vector representations of keywords, but this method comes with notable shortcomings. While word2vec focuses mainly on syntactic relationships, it often falls short in capturing deeper semantic connections. It struggles with words that have multiple meanings and doesn’t thoroughly account for the context in longer pieces of text, which can result in less effective taxonomy structures. In this paper, we propose a novel approach that addresses this issue by utilizing large language models (LLMs) to generate similarity scores between pairs of keywords or key phrases. We then use these similarity scores to train a sentence embedding model that can capture semantic relationships and context-dependent meanings. Unlike word2vec, our approach takes advantage of the detailed contextual embeddings provided by LLMs, leading to more accurate and meaningful taxonomies.