The benefits of conducting a systematic review (SR) within a research project are well recognised. Nonetheless, nowadays an SR demands significant time and effort as it typically requires manually sifting through a large corpus of papers. In this work, we propose an application that simplifies the SR process for researchers by performing topic modeling. Our approach involves more than just conducting topic modeling; it entails selecting the optimal method tailored to our system’s needs. We explore two widely recognized methods in topic modeling, namely Latent Dirichlet Allocation (LDA) and Non-negative Matrix Factorization (NMF), alongside two prominent libraries, Scikit-learn and Gensim, which provide implementations for these methods. The evaluation of these methods and libraries is done with the help of the cosine similarity metric. Our findings indicate that, for our research context, LDA implemented in Scikit-learn stands out as the most effective method for topic modeling. Additionally, our study suggests that employing 10 topics, along with stemming and lemmatization, yields better results. Our findings culminate in the development of an application to facilitate SRs for researchers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Systematic Literature Reviews Methodology Through Topic Modeling

  • Salma Mekaoui,
  • Ilham Chaker,
  • Arsalane Zarghili,
  • Nikola S. Nikolov

摘要

The benefits of conducting a systematic review (SR) within a research project are well recognised. Nonetheless, nowadays an SR demands significant time and effort as it typically requires manually sifting through a large corpus of papers. In this work, we propose an application that simplifies the SR process for researchers by performing topic modeling. Our approach involves more than just conducting topic modeling; it entails selecting the optimal method tailored to our system’s needs. We explore two widely recognized methods in topic modeling, namely Latent Dirichlet Allocation (LDA) and Non-negative Matrix Factorization (NMF), alongside two prominent libraries, Scikit-learn and Gensim, which provide implementations for these methods. The evaluation of these methods and libraries is done with the help of the cosine similarity metric. Our findings indicate that, for our research context, LDA implemented in Scikit-learn stands out as the most effective method for topic modeling. Additionally, our study suggests that employing 10 topics, along with stemming and lemmatization, yields better results. Our findings culminate in the development of an application to facilitate SRs for researchers.