A Blockchain-Based Framework for Crowdsourcing Evaluation of Large Language Models
摘要
The evaluation of outputs from large language models (LLMs) is an important part of LLMs’ born. A comprehensive evaluation for LLMs requires substantial human and material resources. This work proposes a crowdsourcing evaluation framework based on blockchain to comprehensively evaluate the toxicity of the outputs from LLMs. The framework offers LLM service to users and collects their evaluation scores for the outputs from the LLM. The evaluation scores are kept on blockchain. During this, the framework allocates and updates the reputation scores based on users’ contributions to the overall evaluation, in order to mitigate the impact of individual users’ subjective biases on the results. This framework lowers the cost of evaluation and enhances the objectivity and reliability of the results. Experiments demonstrate that this framework provides robust support for the objective evaluation of LLMs and offers a feasible and efficient idea for future works in evaluation for LLMs.