The growing interest in synthetic data generation has led to the development of tabular generative models, which play a critical role in fields like bioinformatics where privacy concerns, data scarcity, and high dimensionality are significant challenges. This study presents a comparative analysis of advanced generative models, including Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and hybrid approaches, focusing on their application to gene-expression datasets. These models are evaluated on their ability to replicate the statistical and biological properties of the data while ensuring utility for predictive tasks. Using rigorous experimental frameworks and real-world datasets, the study explores their strengths, limitations, scalability, and robustness under diverse configurations. Furthermore, it examines the trade-offs between data privacy and utility, shedding light on their suitability for various bioinformatics applications. By identifying key insights and future research directions, this work contributes to the advancement of synthetic biology, facilitating broader adoption of synthetic data methodologies in genomics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Analysis of Tabular Generative Models on Gene-Expression Data

  • Erdenebileg Batbaatar,
  • Keun Ho Ryu

摘要

The growing interest in synthetic data generation has led to the development of tabular generative models, which play a critical role in fields like bioinformatics where privacy concerns, data scarcity, and high dimensionality are significant challenges. This study presents a comparative analysis of advanced generative models, including Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and hybrid approaches, focusing on their application to gene-expression datasets. These models are evaluated on their ability to replicate the statistical and biological properties of the data while ensuring utility for predictive tasks. Using rigorous experimental frameworks and real-world datasets, the study explores their strengths, limitations, scalability, and robustness under diverse configurations. Furthermore, it examines the trade-offs between data privacy and utility, shedding light on their suitability for various bioinformatics applications. By identifying key insights and future research directions, this work contributes to the advancement of synthetic biology, facilitating broader adoption of synthetic data methodologies in genomics.