Bayesian Generation of Synthetic Data
摘要
Generation of synthetic data can be a valuable tool for machine-learning tasks and, in general, managing large volumes of data. This paper presents a technique for creating synthetic data through Bayesian Generation, so that synthetic data maintain the original probability distribution and can be exploited for training Machine-Learning models in place of the original dataset. The paper presents the method and analyzes its impact on selected machine-learning models, by evaluating both the effectiveness and efficiency of the overall process.