Where Are Biases? Adversarial Debiasing with Spurious Feature Visualization
摘要
To avoid deep learning models utilizing shortcuts in a training dataset, many debiasing models have been developed to encourage models learning from accurate correlations. Some research constructs robust models via adversarial training. Although this series of methods shows promising debiasing performance, we do not know precisely what spurious features have been discarded during adversarial training. To address its lack of explainability especially in scenarios with low error tolerance, we design AdvExp, which not only visualizes the underlying spurious feature behind adversarial training but also maintains good debiasing performance with the assistance of a robust optimization algorithm. We show promising performance of AdvExp on BiasCheXpert, a subsampled dataset from CheXpert, and uncover potential regions in radiographs recognized by deep neural networks as gender or race-related features.