<p>Fairness in both machine learning (ML) and human decision-making is crucial, yet both are inherently prone to distinct forms of bias: algorithmic or data-driven in ML models, and subjective or inconsistency-related in humans. This study investigates fairness within the context of university admissions using a real-world dataset of 1,121 applicant profiles, of which 870 correspond to students with local qualifications. We evaluate three ML models, namely the Extreme Gradient Boosting (XGB), Bidirectional Long Short-Term Memory (Bi-LSTM), and k-Nearest Neighbors (KNN), enhanced with BERT embeddings to capture textual features from application materials. To assess <i>individual fairness</i>, we employ a consistency-based metric that measures the agreement between predictions made by ML models and evaluations from human experts of diverse backgrounds. Results demonstrate that ML models exhibit superior consistency compared to human evaluators, with improvements of more than 14%. For <i>group fairness</i>, we propose a gender-debiasing pipeline that systematically reduces explicit gender-indicative linguistic cues while preserving predictive performance. Empirically, debiasing does not degrade classification accuracy for any of the ML models, with point-estimate accuracy improving in some cases, suggesting that fairness and predictive accuracy need not be competing objectives. Overall, our findings demonstrate the potential of hybrid human–ML frameworks to promote more consistent and transparent decision-making processes in high-stakes domains such as university admissions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fairness at no cost: how data debiasing improves equity without compromising accuracy in machine learning models

  • Junhua Liu,
  • Roy Ka-Wei Lee,
  • Kwan Hui Lim

摘要

Fairness in both machine learning (ML) and human decision-making is crucial, yet both are inherently prone to distinct forms of bias: algorithmic or data-driven in ML models, and subjective or inconsistency-related in humans. This study investigates fairness within the context of university admissions using a real-world dataset of 1,121 applicant profiles, of which 870 correspond to students with local qualifications. We evaluate three ML models, namely the Extreme Gradient Boosting (XGB), Bidirectional Long Short-Term Memory (Bi-LSTM), and k-Nearest Neighbors (KNN), enhanced with BERT embeddings to capture textual features from application materials. To assess individual fairness, we employ a consistency-based metric that measures the agreement between predictions made by ML models and evaluations from human experts of diverse backgrounds. Results demonstrate that ML models exhibit superior consistency compared to human evaluators, with improvements of more than 14%. For group fairness, we propose a gender-debiasing pipeline that systematically reduces explicit gender-indicative linguistic cues while preserving predictive performance. Empirically, debiasing does not degrade classification accuracy for any of the ML models, with point-estimate accuracy improving in some cases, suggesting that fairness and predictive accuracy need not be competing objectives. Overall, our findings demonstrate the potential of hybrid human–ML frameworks to promote more consistent and transparent decision-making processes in high-stakes domains such as university admissions.