The rapid advancement of Large Language Models (LLMs) has led to their widespread adoption in various academic and business applications. However, the reliability of these models remains a concern, particularly in situations where their outputs cannot be fully trusted. This paper presents an approach to identify potential errors in large LLM benchmarks by leveraging the consensus of frontier models. Our study focuses on the Massive Multitask Language Understanding (MMLU) benchmark, a popular dataset used to evaluate the performance of LLMs across a wide range of subjects. Our approach demonstrates the potential for using model consensus as a tool to detect benchmark errors and can lead to the creation of cleaner, more accurate datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus

  • Cecilia Delgado-Solorzano,
  • Manuel Delaflor,
  • Carlos Toxtli

摘要

The rapid advancement of Large Language Models (LLMs) has led to their widespread adoption in various academic and business applications. However, the reliability of these models remains a concern, particularly in situations where their outputs cannot be fully trusted. This paper presents an approach to identify potential errors in large LLM benchmarks by leveraging the consensus of frontier models. Our study focuses on the Massive Multitask Language Understanding (MMLU) benchmark, a popular dataset used to evaluate the performance of LLMs across a wide range of subjects. Our approach demonstrates the potential for using model consensus as a tool to detect benchmark errors and can lead to the creation of cleaner, more accurate datasets.