错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing machine learning efficacy and fairness in automated decision systems: an adversarial deep generative modeling with CoBS-TGAN approach in imbalanced and biased datasets

  • Khursheed Ahmad Bhat,
  • Shabir Ahmad Sofi

摘要

Machine learning and deep learning fields have succeeded in their goal of mimicking human behavior and achieving revolutionary performances in data-abundant situations. Still, it is frequently inhibited in data scarcity problems. To tackle this, researchers have come up with the concept of synthetic data to provide a low-cost, easily available, and secure alternative. Synthetic data bolsters the robustness of model learning within real-world contexts, addressing the formidable challenge posed by data scarcity. The constrained availability and scarcity of data result from diverse intrinsic factors, encompassing data regulations, privacy concerns, the confidential nature of data, and the inherent rarity of data of interest in critical real-world applications. This scarcity leads to class imbalance problems often encountered in real world datasets. Popular strategies involve increasing the representation of minority class instances by generating synthetic examples. The existing data augmentation techniques aim to expand datasets for balancing, yet they frequently fail to achieve satisfactory sample diversity. This paper examines the potential of deep learning-based generative adversarial networks for oversampling biased samples. The paper proposes an approach for tabular data involving mixed attributes and pays special attention to data imbalance and improper data labeling. We introduce an adversarial learning based conditioned on biased samples tabular generative adversarial neural network tailored for synthesizing samples on underrepresented target labels and biased attributes in the form of protected or sensitive attributes. This approach exhibits the potential to restore dataset balance, address bias in the dataset, mitigate over-fitting issues, and enhance training data diversity, thereby paying special attention to the downstream classification and generalization performance with fairness considerations. Experiments are conducted on benchmark datasets to validate the feasibility of the model presented in this paper in realistic scenarios. The evaluation and analysis of experimental procedures demonstrate favorable comparisons with other existing synthetic augmentation techniques.