<p>In light of the growing interest in using large language models (LLMs) as tools for generating scientific texts, the evaluation of their ability to produce encyclopedic content is becoming increasingly relevant. However, for Russian-language materials, this issue has not been sufficiently studied and existing benchmarks do not cover key aspects of analytical work with sources. This article presents RuWikiBench—an open benchmark based on Ruwiki for evaluating the ability of LLMs to reproduce Wikipedia-style articles, constructed around three tasks: selection of relevant sources, article structuring, and section generation. The results of testing popular open-source LLMs show that even under ideal conditions, the best models do not always follow the expert logic of composing encyclopedic content: even with a perfect source retrieval system, the models cannot reproduce the reference table of contents, and the quality of section generation shows almost no dependence on the number of parameters.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RuWikiBench: Evaluating Large Language Models Through Replication of Encyclopedia Articles

  • D. A. Grigoriev,
  • D. I. Chernyshev

摘要

In light of the growing interest in using large language models (LLMs) as tools for generating scientific texts, the evaluation of their ability to produce encyclopedic content is becoming increasingly relevant. However, for Russian-language materials, this issue has not been sufficiently studied and existing benchmarks do not cover key aspects of analytical work with sources. This article presents RuWikiBench—an open benchmark based on Ruwiki for evaluating the ability of LLMs to reproduce Wikipedia-style articles, constructed around three tasks: selection of relevant sources, article structuring, and section generation. The results of testing popular open-source LLMs show that even under ideal conditions, the best models do not always follow the expert logic of composing encyclopedic content: even with a perfect source retrieval system, the models cannot reproduce the reference table of contents, and the quality of section generation shows almost no dependence on the number of parameters.