Recently online newspapers and social media flooded information about the dreadful pandemic disease corona and its variants. Corona symptoms, highly affected areas, percentage of corona-affected people, and death rate have been shared by many users on social media. However, due to the swarming of huge data, sequencing and heading vital information is tiresome. This demands analysis of social media data against Covid-19 and subsequent measures to be adhered to. The current study considers the Covid-19 Twitter dataset collected provided by Harvard open-source data. This research works to create Topic modelling using the with Latent Dirichlet Allocation (TMLDA) method for the pandemic COVID-19 Twitter dataset. The collected data is analyzed with Latent Dirichlet Allocation (LDA) and a specific pattern is gathered with the Hclust- agglomerative hierarchical clustering algorithm. The proposed work excels the existing works of Probabilistic Latent Semantic Analysis (PLSA) using Topic sentiment mixture, MaxEnt-LDA and achieved TMLDA 95% of accuracy, 96% of precision, 95% of Recall, and 95% of F1 Score value.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Topic Modelling Using Latent Dirichlet Allocation Model for Pandemic Covid-19 Twitters Dataset

  • K. Sundravadivelu,
  • K. Palaniammal,
  • M. Saravana Roentgen Mani

摘要

Recently online newspapers and social media flooded information about the dreadful pandemic disease corona and its variants. Corona symptoms, highly affected areas, percentage of corona-affected people, and death rate have been shared by many users on social media. However, due to the swarming of huge data, sequencing and heading vital information is tiresome. This demands analysis of social media data against Covid-19 and subsequent measures to be adhered to. The current study considers the Covid-19 Twitter dataset collected provided by Harvard open-source data. This research works to create Topic modelling using the with Latent Dirichlet Allocation (TMLDA) method for the pandemic COVID-19 Twitter dataset. The collected data is analyzed with Latent Dirichlet Allocation (LDA) and a specific pattern is gathered with the Hclust- agglomerative hierarchical clustering algorithm. The proposed work excels the existing works of Probabilistic Latent Semantic Analysis (PLSA) using Topic sentiment mixture, MaxEnt-LDA and achieved TMLDA 95% of accuracy, 96% of precision, 95% of Recall, and 95% of F1 Score value.