<p>High-quality Chinese datasets for public opinion analysis are scarce, with most resources based on outdated English-language platforms like Twitter and Reddit, which are not relevant to Chinese social media. To address this, we introduce CommentAgent, a novel simulation framework that generates high-quality Chinese datasets for sentiment analysis, sarcasm detection, and stance detection. CommentAgent simulates the Weibo environment using a multi-agent architecture that models user behavior, information flow, and personality evolution, allowing for realistic and diverse comment generation. Unlike traditional data augmentation or basic generation methods, CommentAgent generates topic-driven data from scratch. Three benchmark datasets are released for the above tasks, with extensive experiments demonstrating that CommentAgent outperforms existing methods in data quality and generalizability. Moreover, lightweight models trained on these datasets achieve performance comparable to large language models (LLMs). CommentAgent offers a scalable solution for creating up-to-date Chinese datasets for robust public opinion analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CommentAgent: a LLM-powered agent framework for automated comment generation and opinion understanding

  • Xiao Han,
  • Jingyun Sun,
  • Songhua Yu,
  • Libo Qin,
  • Hui Luo,
  • Yang Li

摘要

High-quality Chinese datasets for public opinion analysis are scarce, with most resources based on outdated English-language platforms like Twitter and Reddit, which are not relevant to Chinese social media. To address this, we introduce CommentAgent, a novel simulation framework that generates high-quality Chinese datasets for sentiment analysis, sarcasm detection, and stance detection. CommentAgent simulates the Weibo environment using a multi-agent architecture that models user behavior, information flow, and personality evolution, allowing for realistic and diverse comment generation. Unlike traditional data augmentation or basic generation methods, CommentAgent generates topic-driven data from scratch. Three benchmark datasets are released for the above tasks, with extensive experiments demonstrating that CommentAgent outperforms existing methods in data quality and generalizability. Moreover, lightweight models trained on these datasets achieve performance comparable to large language models (LLMs). CommentAgent offers a scalable solution for creating up-to-date Chinese datasets for robust public opinion analysis.