<p>Topic modeling remains essential for uncovering latent structures in large text corpora, yet performance varies across languages, domains, and document lengths. This study compares five models—Latent Dirichlet Allocation (LDA), Collapsed Gibbs Sampling for LDA, LDA2Vec, Top2Vec, and BERTopic—across three datasets: Hausa news, English short texts (20 Newsgroups), and English long-form corpora (PubMed abstracts and legal case summaries). All models were trained using standardized preprocessing and coherence-based topic optimization under fully reproducible settings. Evaluation combined quantitative metrics (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\({C}_{v}\)</EquationSource> </InlineEquation> coherence and perplexity), expert human assessments (<i>n</i> = 20), and computational profiling. Results show that BERTopic, powered by the multilingual transformer <i>paraphrase-multilingual-MiniLM-L12-v2</i>, achieved the highest coherence (0.67) and interpretability across multilingual and domain-specific corpora but required greater computational resources. LDA with Gibbs Sampling offered a more efficient alternative with competitive coherence on smaller datasets. Overall, the findings reveal a clear trade-off between semantic depth and computational scalability, providing practical guidance for multilingual and domain-specific topic modeling.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Evaluation of Probabilistic and Transformer-Based Topic Models Across Diverse and Multilingual Text Corpora

  • Micheal Olalekan Ajinaja,
  • Johnson Tunde Fakoya,
  • Yetunde Esther Ogunwale,
  • John Kolawole Omoniyi,
  • Michael Adeniyi Ibiyomi,
  • Akeem Adekunle Abiona,
  • Damilola Akinola

摘要

Topic modeling remains essential for uncovering latent structures in large text corpora, yet performance varies across languages, domains, and document lengths. This study compares five models—Latent Dirichlet Allocation (LDA), Collapsed Gibbs Sampling for LDA, LDA2Vec, Top2Vec, and BERTopic—across three datasets: Hausa news, English short texts (20 Newsgroups), and English long-form corpora (PubMed abstracts and legal case summaries). All models were trained using standardized preprocessing and coherence-based topic optimization under fully reproducible settings. Evaluation combined quantitative metrics ( \({C}_{v}\) coherence and perplexity), expert human assessments (n = 20), and computational profiling. Results show that BERTopic, powered by the multilingual transformer paraphrase-multilingual-MiniLM-L12-v2, achieved the highest coherence (0.67) and interpretability across multilingual and domain-specific corpora but required greater computational resources. LDA with Gibbs Sampling offered a more efficient alternative with competitive coherence on smaller datasets. Overall, the findings reveal a clear trade-off between semantic depth and computational scalability, providing practical guidance for multilingual and domain-specific topic modeling.