Are Fair Machine Learning Models More Useful?
摘要
Over the last few years, machine learning (ML) fairness has been widely studied: both its impacts on society, and methods for making it fairer. Yet, interestingly, the relationship between fairness and utility (usefulness) has remained underexplored. The small amount of existing research assumes that fairness and usefulness are conflicting goals, requiring compromise. In contrast, we reason that data and models describing a group of people are often used by those same people, and therefore, models that better describe users will be more useful. In particular, when groups of people are more diverse, the data and models describing them must also be more diverse: fair machine learning models should be more useful (not less). We tested this hypothesis by modeling human color naming—which varies significantly by gender—using datasets with different gender balances and varying fairness. We then compared the model results to actual human naming behavior, finding good evidence that fairer data and models that better represent users are more useful to those users in that they agree with them more often. We conclude with result discussion followed by potential future work.