Ensuring the fairness of machine learning (ML) applications is critical to the reliability of modern artificial intelligence systems. Despite extensive study on this topic, the fairness of ML models in the software engineering (SE) domain has not yet been explored well. As a result, many ML-powered software systems, particularly those utilized in the software engineering community, continue to be prone to fairness issues. Taking one of the typical SE tasks, i.e., code reviewer recommendation, as a subject, this paper investigates the fairness of ML applications in the SE domain, specifically focusing on the code reviewer recommendation task. Our empirical study demonstrates that existing ML-based code reviewer recommendation systems exhibit unfairness and discriminating behaviors. Specifically, male reviewers get, on average, 7.25% more recommendations than female code reviewers compared to their distribution in the reviewer set. This paper also investigates why the studied ML-based code reviewer recommendation systems are unfair and provides solutions to mitigate the unfairness. For instance, such systems may recommend male reviewers at a significantly higher rate than female reviewers in a discriminatory manner. Our study further indicates that the existing mitigation methods can enhance fairness significantly in projects with a similar distribution of protected and privileged groups. Still, their effectiveness in improving fairness on imbalanced or skewed data is limited.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fairness Analysis of Machine Learning-Based Code Reviewer Recommendation

  • Mohammad Mahdi Mohajer,
  • Alvine Boaye Belle,
  • Nima Shiri Harzevili,
  • Junjie Wang,
  • Hadi Hemmati,
  • Song Wang,
  • Zhen Ming (Jack) Jiang

摘要

Ensuring the fairness of machine learning (ML) applications is critical to the reliability of modern artificial intelligence systems. Despite extensive study on this topic, the fairness of ML models in the software engineering (SE) domain has not yet been explored well. As a result, many ML-powered software systems, particularly those utilized in the software engineering community, continue to be prone to fairness issues. Taking one of the typical SE tasks, i.e., code reviewer recommendation, as a subject, this paper investigates the fairness of ML applications in the SE domain, specifically focusing on the code reviewer recommendation task. Our empirical study demonstrates that existing ML-based code reviewer recommendation systems exhibit unfairness and discriminating behaviors. Specifically, male reviewers get, on average, 7.25% more recommendations than female code reviewers compared to their distribution in the reviewer set. This paper also investigates why the studied ML-based code reviewer recommendation systems are unfair and provides solutions to mitigate the unfairness. For instance, such systems may recommend male reviewers at a significantly higher rate than female reviewers in a discriminatory manner. Our study further indicates that the existing mitigation methods can enhance fairness significantly in projects with a similar distribution of protected and privileged groups. Still, their effectiveness in improving fairness on imbalanced or skewed data is limited.