The evaluation of outputs from large language models (LLMs) is an important part of LLMs’ born. A comprehensive evaluation for LLMs requires substantial human and material resources. This work proposes a crowdsourcing evaluation framework based on blockchain to comprehensively evaluate the toxicity of the outputs from LLMs. The framework offers LLM service to users and collects their evaluation scores for the outputs from the LLM. The evaluation scores are kept on blockchain. During this, the framework allocates and updates the reputation scores based on users’ contributions to the overall evaluation, in order to mitigate the impact of individual users’ subjective biases on the results. This framework lowers the cost of evaluation and enhances the objectivity and reliability of the results. Experiments demonstrate that this framework provides robust support for the objective evaluation of LLMs and offers a feasible and efficient idea for future works in evaluation for LLMs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Blockchain-Based Framework for Crowdsourcing Evaluation of Large Language Models

  • Zefeng Mo,
  • Zhihao Hou,
  • Ruilin Lai,
  • Xiaoyuan Wu,
  • Junjie Zhou,
  • Gansen Zhao

摘要

The evaluation of outputs from large language models (LLMs) is an important part of LLMs’ born. A comprehensive evaluation for LLMs requires substantial human and material resources. This work proposes a crowdsourcing evaluation framework based on blockchain to comprehensively evaluate the toxicity of the outputs from LLMs. The framework offers LLM service to users and collects their evaluation scores for the outputs from the LLM. The evaluation scores are kept on blockchain. During this, the framework allocates and updates the reputation scores based on users’ contributions to the overall evaluation, in order to mitigate the impact of individual users’ subjective biases on the results. This framework lowers the cost of evaluation and enhances the objectivity and reliability of the results. Experiments demonstrate that this framework provides robust support for the objective evaluation of LLMs and offers a feasible and efficient idea for future works in evaluation for LLMs.