Topic modelling is an important technique in natural language processing, enabling the extraction of underlying themes from large textual corpora. Measuring the performance of topic models is largely an open problem. In addition, the increasing number of multilingual corpora makes it important to be able to measure the multilingual performance of topic models. Current coherence metrics either fail to take into account multilinguality, or they rely on expensive parallel corpora. To address these gaps, we introduce a novel performance metric called Multilingual Topic Coherence (MTC). MTC measures the extent to which topics represent multilingual ideas without using a parallel corpus. It combines word embedding-based coherence with a multilinguality score. To test MTC’s effectiveness, we constructed a short-text dataset comprising texts from nine of South Africa’s official languages. We compared the MTC performance of three widely used topic models: Latent Dirichlet Allocation (LDA), ZeroShotTM, and BERTopic. Our findings indicate that contextualized topic models, such as ZeroShotTM and BERTopic, outperform traditional models like LDA in terms of both coherence and multilinguality. The MTC metric provides a more nuanced evaluation, highlighting the strengths and weaknesses of each model in handling multilingual topics. This research advances the understanding and application of topic modelling in multilingual and low-resource environments, offering a robust metric for future evaluations. Code used to conduct this study is provided ( https://github.com/AlgorithmicAmoeba/multilingual_topic_coherence ).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing Multilinguality of Topic Models on a Short-Text South African Languages Dataset

  • Darren Craig Roos,
  • Katherine Mary Malan

摘要

Topic modelling is an important technique in natural language processing, enabling the extraction of underlying themes from large textual corpora. Measuring the performance of topic models is largely an open problem. In addition, the increasing number of multilingual corpora makes it important to be able to measure the multilingual performance of topic models. Current coherence metrics either fail to take into account multilinguality, or they rely on expensive parallel corpora. To address these gaps, we introduce a novel performance metric called Multilingual Topic Coherence (MTC). MTC measures the extent to which topics represent multilingual ideas without using a parallel corpus. It combines word embedding-based coherence with a multilinguality score. To test MTC’s effectiveness, we constructed a short-text dataset comprising texts from nine of South Africa’s official languages. We compared the MTC performance of three widely used topic models: Latent Dirichlet Allocation (LDA), ZeroShotTM, and BERTopic. Our findings indicate that contextualized topic models, such as ZeroShotTM and BERTopic, outperform traditional models like LDA in terms of both coherence and multilinguality. The MTC metric provides a more nuanced evaluation, highlighting the strengths and weaknesses of each model in handling multilingual topics. This research advances the understanding and application of topic modelling in multilingual and low-resource environments, offering a robust metric for future evaluations. Code used to conduct this study is provided ( https://github.com/AlgorithmicAmoeba/multilingual_topic_coherence ).