<p>Large Language Models (LLMs) tend to hallucinate when processing symbolically complex linguistic structures. Existing literature evaluates hallucination either through their mechanistic interpretability or at the behavioral output level, but hardly links the symbolic triggers to their layer-wise representational causes. This paper introduces a unified symbolic, behavioral, and mechanistic framework that connects symbolic triggers with internal failure dynamics in transformer architectures. The study evaluates five open-weight LLMs across QA, MCQ, and Odd-One-Out formats on the HaluEval and TruthfulQA datasets, focusing on negation, exceptions, modifiers, numbers, and named entity cues. The results show that hallucination rates remain high across all model scales, with all symbolic categories exhibiting high hallucination rates, and exceptions and numbers often showing comparable or higher values across models. Constrained task formats reduce surface errors but preserve failure patterns, indicating representational instability rather than purely decoding artifacts. Layer-wise analysis shows peak symbolic attention variance in early transformer layers (2–4), after which these patterns persist across deep layers. The consistency of this behavior across architectures suggests that hallucination is strongly associated with weakness in symbolic encoding. The framework provides an interpretable basis for diagnosing and stabilizing symbolic reasoning in LLMs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Layer-wise symbolic attention instability as a diagnostic signal for hallucination in large language models

  • Naveen Lamba,
  • Sanju Tiwari,
  • Manas Gaur

摘要

Large Language Models (LLMs) tend to hallucinate when processing symbolically complex linguistic structures. Existing literature evaluates hallucination either through their mechanistic interpretability or at the behavioral output level, but hardly links the symbolic triggers to their layer-wise representational causes. This paper introduces a unified symbolic, behavioral, and mechanistic framework that connects symbolic triggers with internal failure dynamics in transformer architectures. The study evaluates five open-weight LLMs across QA, MCQ, and Odd-One-Out formats on the HaluEval and TruthfulQA datasets, focusing on negation, exceptions, modifiers, numbers, and named entity cues. The results show that hallucination rates remain high across all model scales, with all symbolic categories exhibiting high hallucination rates, and exceptions and numbers often showing comparable or higher values across models. Constrained task formats reduce surface errors but preserve failure patterns, indicating representational instability rather than purely decoding artifacts. Layer-wise analysis shows peak symbolic attention variance in early transformer layers (2–4), after which these patterns persist across deep layers. The consistency of this behavior across architectures suggests that hallucination is strongly associated with weakness in symbolic encoding. The framework provides an interpretable basis for diagnosing and stabilizing symbolic reasoning in LLMs.