错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated Topic Exploration in a Cultural Heritage Corpus

  • Kyriaki Zoutsou,
  • Michalis Sfakakis,
  • Leonidas Papachristopoulos,
  • Christos Papatheodorou

摘要

The overproduction of scientific literature in cultural heritage has imposed new challenges for the researchers of the field. The emergence of Natural Language Processing methods have increased their ability to unveil latent topical structures in textual corpora, but also demands an analogous effort for both their topic number estimation and interpretation, especially in a nuanced and context-rich content as cultural heritage. This study focuses on the fine-tuning of the most widespread topic modeling algorithm Latent Dirichlet Allocation (LDA), in order to unveil the most representative topics in a corpus of 40881 abstracts of papers dealing with cultural heritage research, retrieved from Scopus. We evaluated the impact of the LDA hyperparameters’ α and β on the coherence and interpretability of the generated topics. Among the systematic experimental attempts, the LDA setting with tuning symmetric α and 0.91 β provided the most coherent topics. Our findings highlight the effectiveness of the LDA algorithm and contribute to the discourse on the need for hyperparameters adjustments to improve the generated results. On the other hand, human interpretation is still needed for the validation of results.