This chapter demonstrates the use of artificial intelligence language models to generate synthetic data for statistical processing and data analysis tasks. The output of working with a large language model (LLM) is the required code in Matlab, which generates the required synthetic data and is part of an automatic generator for parameterized tasks. The output in the form of generating code is preferred over using the LLM online to directly generate synthetic data, which proves to be problematic. This chapter mainly focuses on the formulation of a prompt for the LLM and its modification to achieve synthetic data that most closely approximates real data. The LLM is used to design the attribute structure of tabular data according to the specified domain with the required number of numeric and categorical attributes. Thus, the developed Matlab code is particularly useful in generating real categorical data values and is also able to implement realistic relationships between the values of different categorical attributes. Examples of prompts, their successive modifications, and their outputs are used to show the capabilities of the LLM, as well as their weaknesses and possible solutions to overcome these shortcomings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Using an AI-Based Language Model to Generate Synthetic Statistical Data

  • Mikuláš Gangur,
  • Olga Martinčíková Sojková

摘要

This chapter demonstrates the use of artificial intelligence language models to generate synthetic data for statistical processing and data analysis tasks. The output of working with a large language model (LLM) is the required code in Matlab, which generates the required synthetic data and is part of an automatic generator for parameterized tasks. The output in the form of generating code is preferred over using the LLM online to directly generate synthetic data, which proves to be problematic. This chapter mainly focuses on the formulation of a prompt for the LLM and its modification to achieve synthetic data that most closely approximates real data. The LLM is used to design the attribute structure of tabular data according to the specified domain with the required number of numeric and categorical attributes. Thus, the developed Matlab code is particularly useful in generating real categorical data values and is also able to implement realistic relationships between the values of different categorical attributes. Examples of prompts, their successive modifications, and their outputs are used to show the capabilities of the LLM, as well as their weaknesses and possible solutions to overcome these shortcomings.