Unveiling Indian Public Sentiment: A Topic Modeling Comparison Between LDA and BERTopic on “Mann Ki Baat”
摘要
In recent years, in order to increase transparency and accountability in decision-making and in developing policies that reflect the needs and interests of their citizens, governments recognize the importance of public input. In facilitating this, many governments are providing digital open forums for their citizens to engage in direct dialogue with their elected representatives, allowing for an exchange of ideas and opinions with them. This results in the generation of vast amount of digital text. A manual analysis of such large amounts of text data is not practical. Natural language processing (NLP) techniques are well suited in this context to automatically analyze such text data and uncover the hidden patterns in it. The purpose of the present work is to thoroughly analyze and extract the meaningful insights from the corpus of the English transcripts of “Mann Ki Baat”, an open forum aired by All India Radio and initiated by honourable prime minister of India, Shri. Narendra Modi. Here, in analyzing the text data, the effectiveness of the two most popular topic modelling methods—LDA and BERTopic—was assessed. Overall findings produced by Bertopic demonstrated the strongest performance against LDA.