错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing unsupervised keyphrase extraction through the integration of structural details in embedding-based approaches

  • Ketan Goyal,
  • Saurabh Sharma

摘要

Computational Linguistics or Natural Language Processing (NLP) emerged to enable systems to automatically identify and extract keyphrases from human language texts, to mitigate the exploitation of digital sources. To meet the increasing demand for keyphrase extraction tools, researchers are actively developing new tools that claim to be capable of processing any type of document in any field. In this article, an unsupervised word embedding-based approach for keyphrase extraction is proposed. The proposed method involved enhancing the state-of-the art word embeddings by the use of n-grams. Additionally, the method introduced a unique way to create word vectors by considering significant word vectors and their idf-scores. Our model is able to achieve an F-Score of 0.495 using the combination of Glove, Uni, and Bigrams. The combination of Skip-Gram with Uni and Bigrams also obtained better results for the DUC, SemEval, KDD and Inspec datasets in comparison to state-of-the-art.