The growing reliance on high-quality datasets for artificial intelligence (AI) development highlights the need for synthetic data generation (SDG) to address data scarcity, privacy concerns, and acquisition costs. Large language models (LLMs) have emerged as key tools for SDG, enabling automated synthesis of diverse, high-quality data. Recent advancements have introduced agentic workflows, where multiple LLM-powered agents collaborate to generate high-quality synthetic data. This survey systematically examines architectural approaches in LLM-based SDG, comparing traditional single-LLM methods with agentic workflows. Our analysis reveals that while single-LLM methods are simple to implement, they often require human intervention for curation and filtering. In contrast, agentic workflows improve data quality and diversity but require high computational costs. By maintaining an up-to-date repository [5] of research, tools, and datasets, this work serves as a resource for advancing SDG methodologies and optimizing AI-driven data synthesis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Survey of LLM-Based Methods for Synthetic Data Generation and the Rise of Agentic Workflows

  • Ahmad Alismail,
  • Carsten Lanquillon

摘要

The growing reliance on high-quality datasets for artificial intelligence (AI) development highlights the need for synthetic data generation (SDG) to address data scarcity, privacy concerns, and acquisition costs. Large language models (LLMs) have emerged as key tools for SDG, enabling automated synthesis of diverse, high-quality data. Recent advancements have introduced agentic workflows, where multiple LLM-powered agents collaborate to generate high-quality synthetic data. This survey systematically examines architectural approaches in LLM-based SDG, comparing traditional single-LLM methods with agentic workflows. Our analysis reveals that while single-LLM methods are simple to implement, they often require human intervention for curation and filtering. In contrast, agentic workflows improve data quality and diversity but require high computational costs. By maintaining an up-to-date repository [5] of research, tools, and datasets, this work serves as a resource for advancing SDG methodologies and optimizing AI-driven data synthesis.