Abstract <p>Mechanistic interpretability has identified functional subgraphs within large language models (LLMs), known as transformer circuits (TCs), that appear to implement specific algorithms. Yet we lack a formal, single-pass way to quantify when an active circuit is behaving coherently and thus likely trustworthy. Building on the author’s prior sheaf-theoretic formulation of causal emergence (Krasnovsky, 2025) we specialize it to transformer circuits and introduce the single-pass, dimensionless effective-information consistency score (EICS). EICS combines a normalized sheaf inconsistency computed from local Jacobians and activations with a Gaussian EI proxy for circuit-level causal emergence derived from the same forward state. The construction is white-box, single-pass, and makes units explicit so that the score is dimensionless. We further provide practical guidance on score interpretation, computational overhead (with fast and exact modes), and a toy sanity-check analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Measuring Uncertainty in Transformer Circuits with Effective Information Consistency

  • A. A. Krasnovsky

摘要

Abstract

Mechanistic interpretability has identified functional subgraphs within large language models (LLMs), known as transformer circuits (TCs), that appear to implement specific algorithms. Yet we lack a formal, single-pass way to quantify when an active circuit is behaving coherently and thus likely trustworthy. Building on the author’s prior sheaf-theoretic formulation of causal emergence (Krasnovsky, 2025) we specialize it to transformer circuits and introduce the single-pass, dimensionless effective-information consistency score (EICS). EICS combines a normalized sheaf inconsistency computed from local Jacobians and activations with a Gaussian EI proxy for circuit-level causal emergence derived from the same forward state. The construction is white-box, single-pass, and makes units explicit so that the score is dimensionless. We further provide practical guidance on score interpretation, computational overhead (with fast and exact modes), and a toy sanity-check analysis.