错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality

  • Jussi Karlgren,
  • Luise Dürlich,
  • Evangelia Gogoulou,
  • Liane Guillou,
  • Joakim Nivre,
  • Magnus Sahlgren,
  • Aarne Talman

摘要

ELOQUENT is a set of shared tasks for evaluating the quality and usefulness of generative language models. ELOQUENT aims to bring together some high-level quality criteria, grounded in experiences from deploying models in real-life tasks, and to formulate tests for those criteria, preferably implemented to require minimal human assessment effort and in a multilingual setting. The selected tasks for this first year of ELOQUENT are (1) probing a language model for topical competence; (2) assessing the ability of models to generate and detect hallucinations; (3) assessing the robustness of a model output given variation in the input prompts; and (4) establishing the possibility to distinguish human-generated text from machine-generated text.