This study compares the quality of synthetic data generated by four generative AI algorithms: Conditional Generative Adversarial Networks (CGAN), Variational Autoencoders (VAE), Synthetic Minority Over-sampling Technique for Regression (SMOTER), and Time Generative Adversarial Networks (TimeGAN). These models are trained on groundwater well data sourced from the Hanford site in Washington, USA, which has substantial missing data samples. Using these models, additional data were generated to replace these missing values. The datasets generated from each model are evaluated using a Gradient Boosting algorithm, and their data distributions are visualized. Comparative analysis reveals that the CGAN model produced synthetic data with a probability distribution closest to the original dataset. Furthermore, when the CGAN-generated dataset was combined with the original dataset, it yielded the best performance metrics among all generated datasets, demonstrating its excellence in effectively generating high-quality synthetic data. This highlights the potential of CGANs in addressing the challenges posed by missing data in environmental studies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generative AI Techniques for the Simulation of Groundwater Well Data at Hanford Site

  • Fabiola Rivera-Noriega,
  • Alejandro De La Noval,
  • Himanshu Upadhyay,
  • Leonel Lagos,
  • Jayesh Soni

摘要

This study compares the quality of synthetic data generated by four generative AI algorithms: Conditional Generative Adversarial Networks (CGAN), Variational Autoencoders (VAE), Synthetic Minority Over-sampling Technique for Regression (SMOTER), and Time Generative Adversarial Networks (TimeGAN). These models are trained on groundwater well data sourced from the Hanford site in Washington, USA, which has substantial missing data samples. Using these models, additional data were generated to replace these missing values. The datasets generated from each model are evaluated using a Gradient Boosting algorithm, and their data distributions are visualized. Comparative analysis reveals that the CGAN model produced synthetic data with a probability distribution closest to the original dataset. Furthermore, when the CGAN-generated dataset was combined with the original dataset, it yielded the best performance metrics among all generated datasets, demonstrating its excellence in effectively generating high-quality synthetic data. This highlights the potential of CGANs in addressing the challenges posed by missing data in environmental studies.