Comparative Analysis of LDA Algorithm for Low Resource Indian Languages with Its Translated English Documents
摘要
Nowadays, social media acts as an information medium that generates a huge amount of data. Irrespective of the age groups, many people across the globe use social media to share their thoughts and emotions in the form of tweets, comments, etc. This information can be used as a source of data for research. This generates a humongous amount of data not only in English but also in many other languages. In this paper, we have focused our research on native Indian languages such as Kannada, Tamil, and Telugu. Latent Dirichlet Allocation (LDA) is an algorithm that can process a large set of text data to produce topic modeling, clustering words, and topics in a document. We have analyzed the performance of the LDA method for native language. Dataset in Kannada and Tamil works better with respect to coherence and English translated dataset of Telugu is Optimal when compared with original Telugu data. With respect to perplexity, LDA works better for native language dataset.