错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Parity benchmark for measuring bias in LLMs

  • Shmona Simpson,
  • Jonathan Nukpezah,
  • Kie Brooks,
  • Raaghav Pandya

摘要

Bias in Large Language Models (LLMs) can perpetuate harmful stereotypes, reinforce inequities, and lead to unfair outcomes in applications from automated content moderation to decision-making systems. These biases also limit the applicability of LLMs in areas such as law, medicine, education, and finance. This paper introduces a benchmark designed to measure and evaluate biases in LLMs. It addresses the protected characteristics on which bias is often enacted, including gender, race, socioeconomic status, and intersectional identities. By systematically assessing LLMs using an expert-curated dataset, the benchmark tests for the biases present in recent large language models like GPT-4o, Llama 3, Gemini and Claude 3.5 Sonnet. This paper details the construction of the benchmark, including the selection of the categories (Ageism, Colonial bias, Colorism, Disability, Homophobia, Racism, Sexism, and Supremacism), the evaluation metrics, and the implementation of testing protocols. Through empirical analysis, we evaluated the LLMs and observed significant performance disparities in multiple categories. All LLMs had an accuracy of at least 74% on average when tested for knowledge regarding these categories. However, this threshold was reduced when LLMs were required to interpret, reason, or deduce. This was especially true regarding homophobia, colonial praxis, and disability. GPT-4 performed best regarding content knowledge followed closely by Claude 3.5 Sonnet, while Gemma-1.1 performed best with interpretation. Gemini 1.5 Pro was better overall than its predecessor, Gemini 1.0, demonstrating that rapid improvement in bias mitigation is possible. These findings highlight the critical need for ongoing monitoring and mitigation strategies to address bias in generative AI systems. We hope that it will serve as a critical tool for policymakers aiming to promote fairness and provide an opportunity for LLM developers to leverage the benchmark to unlock use cases where they may better serve all of us.