In natural language processing (NLP), hallucinations–factual inaccuracies in model outputs—pose serious challenges, particularly in text summarization. Traditional evaluation methods often struggle to assess the factual consistency of summaries. To address this, we present FARE, a novel multi-step evaluation framework that leverages coreference resolution, atomic fact generation, and vector alignment to evaluate factual consistency. Unlike prior methods that align across entire passages, FARE’s two-stage generation and alignment strategy greatly improves accuracy, interpretability, and efficiency in aligning facts with source content. The experimental results on the AGGREFACT benchmarks show that FARE outperforms existing approaches on several datasets, marking an advance in reliable content evaluation for NLP-generated summaries.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FARE: Factual Alignment for Reliable Evaluation in Text Summarization

  • Zhijian Wang,
  • Cangqi Zhou,
  • Jing Zhang,
  • Dianming Hu

摘要

In natural language processing (NLP), hallucinations–factual inaccuracies in model outputs—pose serious challenges, particularly in text summarization. Traditional evaluation methods often struggle to assess the factual consistency of summaries. To address this, we present FARE, a novel multi-step evaluation framework that leverages coreference resolution, atomic fact generation, and vector alignment to evaluate factual consistency. Unlike prior methods that align across entire passages, FARE’s two-stage generation and alignment strategy greatly improves accuracy, interpretability, and efficiency in aligning facts with source content. The experimental results on the AGGREFACT benchmarks show that FARE outperforms existing approaches on several datasets, marking an advance in reliable content evaluation for NLP-generated summaries.