The Advent of Topic-Noise Models
摘要
As data sets grow in size, determining the content of different types of shared text continues to be a difficult task. The task is even more difficult when the documents are short and noisy. Examples include social media posts, product reviews, clinician notes, and open-ended survey responses. This chapter discusses the emergence of a new class of topic models, topic-noise models. Topic-noise models are generative and treat topics and noise as separate distributions over words. In other words, the model assumes that noise exists and it must be modeled as well. The chapter begins by defining and presenting different topic-noise algorithms, both unsupervised and semi-supervised. We then present the similarities and differences between topic-noise models and well-known topic models like LDA, highlighting when each model performs best. Finally, we consider how these models can be advanced by employing large language models and other auxiliary data sources. In an era when noise can no longer be ignored, topic-noise models offer an important alternative to traditional topic models. They may eventually be the next step in the evolution of topic models.