An Empirical Study on the Impact of Proxy Attributes on the Fairness of Machine Learning Systems
摘要
Fairness within machine learning (ML) has been a subject of considerable interest in recent years. The outcome fairness is usually measured along the protected or the sensitive attributes. However, it remains unclear whether including protected attributes in the feature set alone causes such biases. This study aims to provide significant empirical evidence concerning the existence of proxy attributes and their implications for machine learning outcomes. To this end, two well-recognized fairness datasets were utilized, and experiments were conducted to measure fairness metrics to measure gender bias by constructing machine learning models using three distinct algorithms. Fairness metrics were employed to evaluate the disparity in the outcomes produced by the models constructed with and without gender as a feature for both datasets under examination. The results demonstrate that the fairness metric mainly remained consistent regardless of the inclusion of the gender attribute in the model.