Malaysia Rice Yield Prediction Using Synthetic Data Generation of CTAB-GAN
摘要
Machine learning and artificial intelligence advancement have significantly impacted various sectors, including agriculture. One prominent challenge researchers face in this field is reliable and accurate data availability. This paper explores the use of conditional tabular GAN (CTAB-GAN) to generate synthetic data to enhance agricultural yield prediction models, focusing on rice yield in Malaysia. The approach involves data collection, preparation for analysis, synthetic data generation at various scales, and model training. The study compares synthetic datasets of different sizes—25%, 50%, 75%, 100%, and an additional set of 250 rows—against the actual dataset. The results indicate that while synthetic data generally led to a performance decline in most models, decision tree-based models and the XGBoost Regressor improved performance with smaller synthetic datasets (25% of the original data). Notable improvements were observed in model performance when evaluated using four standard regression metrics. The findings suggest that CTAB-GAN can enhance model performance in specific agricultural contexts by generating synthetic data. It demonstrates how synthetic data can help resolve the problem of insufficient data, improve the precision of crop yield predictions, and contribute to more effective agricultural practices.