错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

REM: A Ranking-Based Automatic Evaluation Method for LLMs

  • Jintao Yang,
  • Yushan Tan,
  • Wenpeng Hu,
  • Zonghao Yang,
  • Xian Zhou,
  • Zhunchen Luo,
  • Wei Luo

摘要

The emergence of Large Language Models (LLMs) has garnered attention due to their remarkable comprehension and generation capabilities across various language tasks and application scenarios. However, traditional evaluation methods have exhibited limitations in terms of fairness and bias. Consequently, there is a pressing need for new evaluation techniques to accurately assess the performance of LLMs. This paper introduces a Ranking-based automatic Evaluation Method (REM) for LLMs. The notable advantage of REM lies in its impartiality, as it harnesses the collective intelligence of multiple LLMs through a voting mechanism, thereby estimating model performance and rankings. The effectiveness of our approach is theoretically supported by numerical estimation, further bolstering its credibility. Subsequent practical tests were carried out on seven open-source models using Chinese and English evaluation datasets, confirming the efficacy of this method in real-world scenarios. By introducing an unbiased automatic evaluation method, this paper makes a valuable contribution to the evaluation of LLMs, eliminating the need for manual intervention by individuals.