Large Language Models (LLMs) exhibit significant capabilities and are extensively applied in diverse critical domains: advanced question-answering systems, healthcare information analysis, etc. Nevertheless, LLMs are prone to hallucination, which refers to generating outputs that are factually incorrect or unsubstantiated. This issue engenders significant reliability risks in critical applications. Inference-time activation editing has emerged as a promising strategy to mitigate hallucination without the need for model retraining. However, existing methods often employ generalized criteria for selecting attention heads and apply editing strengths that lack query-specific adaptability, therefore leading to suboptimal hallucination correction. To address these limitations, we introduce Query-adaptive Saliency-localized Activation Editing (QSE), which comprises Gradient-guided Head Saliency Localization (GSL) and Query-specific Editing Necessity Estimation (QNE), to enhance the precision and contextual adaptability of LLM activation editing. Specifically, GSL first employs a gradient-based optimization process to quantify the differential saliency of attention heads concerning factual generation, thereby pinpointing critical attention heads for precise activation editing. Subsequently, QNE comprehensively perceives the input query’s knowledge semantics, and its lightweight estimator dynamically adjusts the editing strength for each head previously identified by GSL, thereby enabling highly adaptive and context-aware adjustments. Empirical evaluations on the LLaMA-3-8B-Instruct model using the TruthfulQA benchmark demonstrate that QSE achieves substantial improvements in model truthfulness, notably surpassing the baseline by 21.3% on the True*Info score.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

QSE: Mitigating LLM Hallucinations Through Query-Adaptive Saliency-Localized Activation Editing

  • Kewei Liao,
  • Tianbo Wang,
  • Fengxiang Yu

摘要

Large Language Models (LLMs) exhibit significant capabilities and are extensively applied in diverse critical domains: advanced question-answering systems, healthcare information analysis, etc. Nevertheless, LLMs are prone to hallucination, which refers to generating outputs that are factually incorrect or unsubstantiated. This issue engenders significant reliability risks in critical applications. Inference-time activation editing has emerged as a promising strategy to mitigate hallucination without the need for model retraining. However, existing methods often employ generalized criteria for selecting attention heads and apply editing strengths that lack query-specific adaptability, therefore leading to suboptimal hallucination correction. To address these limitations, we introduce Query-adaptive Saliency-localized Activation Editing (QSE), which comprises Gradient-guided Head Saliency Localization (GSL) and Query-specific Editing Necessity Estimation (QNE), to enhance the precision and contextual adaptability of LLM activation editing. Specifically, GSL first employs a gradient-based optimization process to quantify the differential saliency of attention heads concerning factual generation, thereby pinpointing critical attention heads for precise activation editing. Subsequently, QNE comprehensively perceives the input query’s knowledge semantics, and its lightweight estimator dynamically adjusts the editing strength for each head previously identified by GSL, thereby enabling highly adaptive and context-aware adjustments. Empirical evaluations on the LLaMA-3-8B-Instruct model using the TruthfulQA benchmark demonstrate that QSE achieves substantial improvements in model truthfulness, notably surpassing the baseline by 21.3% on the True*Info score.