As dialogue systems and chatbots become more common in daily life, efficient and accurate evaluation methods are crucial. This study compares human and AI assessments across various dialogue scenarios, focusing on seven key performance indicators (KPIs): Coherence, Innovation, Concreteness, Goal Contribution, Commonsense Contradiction, Incorrect Fact, and Redundancy. Using the GPT-4o API, we generated diverse dialogue datasets and conducted a two-part analysis. Experiment 1 evaluated multi-party dialogues on Coherence, Innovation, Concreteness, and Goal Contribution, finding that GPT models closely match human judgments. Experiment 2 focused on dyadic dialogues, assessing Commonsense Contradiction, Incorrect Fact, and Redundancy, showing GPT-4o excels in maintaining factual accuracy and commonsense reasoning but struggles with redundancy and self-contradiction. These findings highlight GPT models’ potential to replicate human evaluation in dialogue systems while identifying areas for improvement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automating Dialogue Evaluation: LLMs Vs Human Judgment

  • Ebubechukwu Ike,
  • Johane Takeuchi,
  • Frank Joublin,
  • Antonello Ceravola,
  • Marc Tanti

摘要

As dialogue systems and chatbots become more common in daily life, efficient and accurate evaluation methods are crucial. This study compares human and AI assessments across various dialogue scenarios, focusing on seven key performance indicators (KPIs): Coherence, Innovation, Concreteness, Goal Contribution, Commonsense Contradiction, Incorrect Fact, and Redundancy. Using the GPT-4o API, we generated diverse dialogue datasets and conducted a two-part analysis. Experiment 1 evaluated multi-party dialogues on Coherence, Innovation, Concreteness, and Goal Contribution, finding that GPT models closely match human judgments. Experiment 2 focused on dyadic dialogues, assessing Commonsense Contradiction, Incorrect Fact, and Redundancy, showing GPT-4o excels in maintaining factual accuracy and commonsense reasoning but struggles with redundancy and self-contradiction. These findings highlight GPT models’ potential to replicate human evaluation in dialogue systems while identifying areas for improvement.