Interpretable White-Box Fairness Testing Through Biased Neuron Identification
摘要
In recent years, deep neural networks (DNNs) have been widely used in a wide range of applications. However, there is a societal concern about the ability of DNNs to make sound and equitable decisions, particularly when they are used in sensitive areas where valuable resources are allocated, such as education, loan, and employment. Before reliable deployment of DNNs in such a sensitive domain, it is essential to do a fair test, i.e., generating as many instances as possible to uncover fairness violations. However, the current testing methods are still restricted in the aspects of interpretability, performance, and generalizability. To overcome the challenges, we propose a new DNN fairness testing framework that differs from previous work in several key aspects: (1) interpretable—it quantitatively interprets DNNs’ fairness violations for the biased decision; (2) effective—it uses the interpretation results to guide the generation of more diverse instances in less time; (3) generic—it can handle both structured and unstructured data. A large number of DNNs are used to evaluate the performance of our method. For example, on a structured dataset, it is also possible to exploit the instances of our method to increase the fairness of the biased DNNs.