错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Unsupervised Keyword Extraction Algorithms from Online Senegalese News Articles

  • Tony Tona Landu,
  • Mamadou Bousso,
  • Mor Absa Loum,
  • Ibrahim Sawadogo,
  • Yoro Dia,
  • Ousmane Sall,
  • Lamine Faty,
  • Ramiyou Karim Mache,
  • Mohamed Sylla

摘要

The production of news article in the online press is growing rapidly. It is clear that these social media sites are filled with more data, which can also mean more difficulties for humans to quickly retrieve or extract specific information according to their needs. Thus, it is more convenient to use automatic keywords, which can be seen as an unsupervised learning problem. Here, we consider a text document as a set of words from the document content. To face this challenge, in this paper, several keyword algorithms, namely KeyBERT, Yake, Rake, etc., have been exploited, and evaluated. The performance of each algorithm was compared with a dataset of news articles written in French language collected from different online news sites in Senegal. The overall ranking result obtained from the experiment using the basic state-of-the-art evaluation methods shows that the Rake extractor achieved a significant improvement in terms of efficiency and time when extracting keywords with 88% of the keywords that most closely match human readers and took 142 seconds when extracting the keywords from the set of documents in our corpus used. Moreover, with our best algorithm presented, journalists or media professionals can use it to improve their work process. This could lead to the emergence of online journalism. Moreover, the list of keywords obtained from our best model can be used for a variety of applications such as document indexing.