As a result of the research done in scientific studies on every subject and the advancement of technology, the number of online research journals has increased, resulting in a deep research pool that cannot be thoroughly examined. As a result of this information pool’s size, it becomes difficult to reach the desired information, and the time spent is considerably increasing. To solve all these problems, it is essential to create Natural Language Processing techniques that will facilitate the examination, narrow the pool, and provide more accessible results. Latent Dirichlet Allocation is a popular Natural Language Processing algorithm for finding hidden topics in long texts. Using the Latent Dirichlet Allocation method, this study aims to examine the research in the journals and determine the subjects covered by the journal. For this goal, articles were gathered from many journals. The gathered articles’ texts were cleaned through various steps. Bigrams and trigrams were obtained after creating a token list. After making the corpus, a topic identification study was carried out with the LDA model for 30 topics. As a result of the study, different parameters are effective in Latent Dirichlet Allocation applications. Selecting appropriate Latent Dirichlet Allocation parameters may achieve more meaningful and accurate results. It is seen that various pre-processes need to be applied, and it has been determined that a text in a different language affects the performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Topic Modeling on Journals with Latent Dirichlet Allocation

  • Mert Çakar,
  • Yusuf Kartal,
  • Eyyüp Gülbandılar

摘要

As a result of the research done in scientific studies on every subject and the advancement of technology, the number of online research journals has increased, resulting in a deep research pool that cannot be thoroughly examined. As a result of this information pool’s size, it becomes difficult to reach the desired information, and the time spent is considerably increasing. To solve all these problems, it is essential to create Natural Language Processing techniques that will facilitate the examination, narrow the pool, and provide more accessible results. Latent Dirichlet Allocation is a popular Natural Language Processing algorithm for finding hidden topics in long texts. Using the Latent Dirichlet Allocation method, this study aims to examine the research in the journals and determine the subjects covered by the journal. For this goal, articles were gathered from many journals. The gathered articles’ texts were cleaned through various steps. Bigrams and trigrams were obtained after creating a token list. After making the corpus, a topic identification study was carried out with the LDA model for 30 topics. As a result of the study, different parameters are effective in Latent Dirichlet Allocation applications. Selecting appropriate Latent Dirichlet Allocation parameters may achieve more meaningful and accurate results. It is seen that various pre-processes need to be applied, and it has been determined that a text in a different language affects the performance.