<p>With the acceleration of Large Language Model (LLM) use by the public, there is an urgent need to make sure the downstream effects have beneficial impacts on humans and society. Therefore, there is an increasing push for ethical evaluation of LLMs to look for potential bias, toxic behaviour, and misinformation. Thus, crowdsourcing has become a popular practice, such as OpenAI’s 2023 Evals initiative. Firstly, by reviewing the literature from software development and ethics, we wish to highlight several cautions on applying the crowdsourcing model to LLMs: including participant self-selection and non-representativeness; the diffusion of responsibility effect including ethics washing and burden shifting; and requisition of ‘incentives’ <i>vis-a-vis</i> issues faced by gig workers. Using the Evals GitHub repository as a case study, we study the effectiveness of an expert-driven, voluntary, crowdsourced scheme on GitHub to address socioethical issues in LLMs. This is achieved by evaluating the statistics of crowdsourced contributions on ethical and bias considerations, which pales in comparison to other technical contributions. This commentary hopes to highlight the issues of ethics, equity, and justice in LLM crowdsourcing, drawing upon interdisciplinary literature, and presents open considerations on how we can improve the state of play.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

“Lost in the crowd”: ethical concerns in crowdsourced evaluations of LLMs

  • Marc Cheong,
  • Gabby Bush,
  • Michael Wildenauer

摘要

With the acceleration of Large Language Model (LLM) use by the public, there is an urgent need to make sure the downstream effects have beneficial impacts on humans and society. Therefore, there is an increasing push for ethical evaluation of LLMs to look for potential bias, toxic behaviour, and misinformation. Thus, crowdsourcing has become a popular practice, such as OpenAI’s 2023 Evals initiative. Firstly, by reviewing the literature from software development and ethics, we wish to highlight several cautions on applying the crowdsourcing model to LLMs: including participant self-selection and non-representativeness; the diffusion of responsibility effect including ethics washing and burden shifting; and requisition of ‘incentives’ vis-a-vis issues faced by gig workers. Using the Evals GitHub repository as a case study, we study the effectiveness of an expert-driven, voluntary, crowdsourced scheme on GitHub to address socioethical issues in LLMs. This is achieved by evaluating the statistics of crowdsourced contributions on ethical and bias considerations, which pales in comparison to other technical contributions. This commentary hopes to highlight the issues of ethics, equity, and justice in LLM crowdsourcing, drawing upon interdisciplinary literature, and presents open considerations on how we can improve the state of play.