SD-WEAT: Towards Robustly Measuring Bias in Input Embeddings
摘要
Artificial intelligence (AI) is rapidly being adopted to build products and aid in the decision-making process across industries. However, AI systems have been shown to exhibit and even amplify biases, causing a growing concern among people worldwide. Thus, investigating methods of measuring and mitigating bias within these AI-powered tools is necessary. In this study, we introduce SD-WEAT, which is a modified version of the Word Embedding Association Test (WEAT) that utilizes the standard deviation (SD) of multiple permutations of the WEAT benchmarks in order to calculate bias in input embeddings, a common area of measuring and mitigating bias in AI. This method produces results comparable to that of WEAT, while addressing some of its largest limitations. Thus, SD-WEAT shows promise for robustly measuring bias in the input embeddings fed to AI language models. Moreover, using approaches from the field Human Computer Interaction (HCI), SD-WEAT is a more accessible and user-friendly method of measuring bias in input embeddings.