错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Artificial Intelligence Models Using Synthetic Data

  • Marina Soledad Iantorno,
  • James Garza

摘要

In the rapidly evolving digital era, Machine Learning (ML) presents a constant need for large quantities and, simultaneously, high-quality data. This research assesses synthetic data and its critical role in ML. The analysis delves into the benefits of using synthetic data instead of real-world datasets. In an era defined by an unending flood of data and privacy regulations becoming more rigorous, synthetic data represents a groundbreaking response to the issues confronting technical professionals and decision-makers [1]. The research employs datasets from the UCI Machine Learning Repository and Kaggle, synthesising data through two different methods: Wasserstein Generative Adversarial Networks (WGANs) and Boolean. These techniques maintain the attributes and patterns of the original datasets while offering diverse advantages. The analysis focuses on evaluating the effectiveness of synthetic data in training ML models, its role in optimising resource allocation, and reducing the time and computational costs associated with collecting real data. Furthermore, the study addresses the risks and ethical concerns related to synthetic data, emphasising the need for responsible standards in its creation and use. This study aims to explain the relevance of synthetic data, its characteristics, and applications and provide insights into its potential to transform the landscape of ML models in the world of data privacy.