FARE: Factual Alignment for Reliable Evaluation in Text Summarization
摘要
In natural language processing (NLP), hallucinations–factual inaccuracies in model outputs—pose serious challenges, particularly in text summarization. Traditional evaluation methods often struggle to assess the factual consistency of summaries. To address this, we present FARE, a novel multi-step evaluation framework that leverages coreference resolution, atomic fact generation, and vector alignment to evaluate factual consistency. Unlike prior methods that align across entire passages, FARE’s two-stage generation and alignment strategy greatly improves accuracy, interpretability, and efficiency in aligning facts with source content. The experimental results on the AGGREFACT benchmarks show that FARE outperforms existing approaches on several datasets, marking an advance in reliable content evaluation for NLP-generated summaries.