Automating Dialogue Evaluation: LLMs Vs Human Judgment
摘要
As dialogue systems and chatbots become more common in daily life, efficient and accurate evaluation methods are crucial. This study compares human and AI assessments across various dialogue scenarios, focusing on seven key performance indicators (KPIs): Coherence, Innovation, Concreteness, Goal Contribution, Commonsense Contradiction, Incorrect Fact, and Redundancy. Using the GPT-4o API, we generated diverse dialogue datasets and conducted a two-part analysis. Experiment 1 evaluated multi-party dialogues on Coherence, Innovation, Concreteness, and Goal Contribution, finding that GPT models closely match human judgments. Experiment 2 focused on dyadic dialogues, assessing Commonsense Contradiction, Incorrect Fact, and Redundancy, showing GPT-4o excels in maintaining factual accuracy and commonsense reasoning but struggles with redundancy and self-contradiction. These findings highlight GPT models’ potential to replicate human evaluation in dialogue systems while identifying areas for improvement.