Techniques to Improve the Utility of Synthetic Data Using XAI and Tabular GAN-Based Approaches
摘要
High quality data can be valuable information for firms, public officials and researchers. Ensuring data privacy without sacrificing quality when dealing with sensitive data is crucial. The need for privacy-preserving and high-quality data has become an important topic in the era of big data. With this background, interest in generating of synthetic data has been increasing in the financial industry. Synthetic data defines as data generated using a purpose-built mathematical model or algorithm with the aim of solving data science task(s). The enhancement of synthetic sample quality can be challenging. Revising the architecture of generative models may not always be effective, depending on the nature of the data. In general, further processing of analytical datasets is more effective. This paper explains how certain features of tabular data do not work well in the existing generative models using XAI. Furthermore, it provides solutions for postprocessing techniques that improve the utility of existing tabular-based generation models by performing default loan prediction and proving that help to improve ML utility.