Twitter and other social media platforms have become important sources of user-generated content, providing vast databases of human activities and opinions for understanding and segmenting users based on behavior, preferences, and demographics is critical for applications including personalized marketing, recommendations, and sentiment analysis. This study explores the feasibility of combining Dirichlet latent classification (LDA) and Bidirectional Encoder Representations from Transformers (BERT) interpolation to improve topic modeling on textual data sets. Traditional LDA and BERT-based clustering have their individual strengths and limitations. LDA excels at generating interpretable topics, while BERT embeddings provide a finer sense of context. We propose a hybrid method, LDA_BERT, which takes advantage of the contextual depth of LDA definition and BERT embedded to improve subject matching and clustering quality. This approach involves first processing a large dataset of textual data, creating BERT embeddings, clustering these embeddings using K-Means, and then applying LDA to the clustered data. We evaluate models using coherence and silhouette scores consider. Our findings show that the LDA_BERT hybrid model outperforms both traditional LDA and BERT-based clustering, with coherence scores improving from 0.24 (LDA) and 0.55 (BERT) to 0.65 (LDA_BERT) and silhouette scores 0.03 (LDA) and 0.21 (BERT) to 0.30 (LDA_BERT). These results indicate that the hybrid approach effectively combines the strengths of both methods, producing consistent and well-defined products. The results of this study are important for areas where complex topic modeling is needed such as social media analytics, recommendation systems, natural language understanding. Future work will focus on the hybrid model further and explore its application use data types and languages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

User Classification on Twitter Using Hybrid Modeling with BERT and LDA

  • Rachna Narula,
  • Md. Sehr Shamir,
  • Nitish Sharma,
  • Paramveer Singh,
  • Anju,
  • Vijay Kumar

摘要

Twitter and other social media platforms have become important sources of user-generated content, providing vast databases of human activities and opinions for understanding and segmenting users based on behavior, preferences, and demographics is critical for applications including personalized marketing, recommendations, and sentiment analysis. This study explores the feasibility of combining Dirichlet latent classification (LDA) and Bidirectional Encoder Representations from Transformers (BERT) interpolation to improve topic modeling on textual data sets. Traditional LDA and BERT-based clustering have their individual strengths and limitations. LDA excels at generating interpretable topics, while BERT embeddings provide a finer sense of context. We propose a hybrid method, LDA_BERT, which takes advantage of the contextual depth of LDA definition and BERT embedded to improve subject matching and clustering quality. This approach involves first processing a large dataset of textual data, creating BERT embeddings, clustering these embeddings using K-Means, and then applying LDA to the clustered data. We evaluate models using coherence and silhouette scores consider. Our findings show that the LDA_BERT hybrid model outperforms both traditional LDA and BERT-based clustering, with coherence scores improving from 0.24 (LDA) and 0.55 (BERT) to 0.65 (LDA_BERT) and silhouette scores 0.03 (LDA) and 0.21 (BERT) to 0.30 (LDA_BERT). These results indicate that the hybrid approach effectively combines the strengths of both methods, producing consistent and well-defined products. The results of this study are important for areas where complex topic modeling is needed such as social media analytics, recommendation systems, natural language understanding. Future work will focus on the hybrid model further and explore its application use data types and languages.