Log parsing transforms unstructured log messages into structured formats and serves as a critical role for a wide range of downstream log analysis tasks. While recent advances in large language models (LLMs) have demonstrated strong performance in log parsing, they tend to rely on spurious correlations and heuristic cues, which introduce bias and compromise parsing accuracy. To address this issue, we propose DebiasParser, a novel log parsing framework that incorporates a debiasing mechanism grounded in Structural Causal Models (SCMs) and implemented via front-door adjustment. Our method does not require access to LLM internals; instead, it leverages carefully designed prompt engineering to model the causal effect between log messages and their templates. To further refine the causal estimation, we incorporate counterfactual rewriting and retrieval-augmented generation (RAG) to enable accurate and unbiased inference over mediators. Empirical evaluation on multiple public log benchmark datasets demonstrates that DebiasParser achieves state-of-the-art performance, significantly improving both log grouping and template extraction accuracy. These findings validate the effectiveness of causal debiasing in enhancing LLM-based log parsing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DebiasParser: Debiasing LLM-Based Log Parsing via Front-Door Adjustment

  • Yuan Tian,
  • Shi Ying,
  • Tiangang Li

摘要

Log parsing transforms unstructured log messages into structured formats and serves as a critical role for a wide range of downstream log analysis tasks. While recent advances in large language models (LLMs) have demonstrated strong performance in log parsing, they tend to rely on spurious correlations and heuristic cues, which introduce bias and compromise parsing accuracy. To address this issue, we propose DebiasParser, a novel log parsing framework that incorporates a debiasing mechanism grounded in Structural Causal Models (SCMs) and implemented via front-door adjustment. Our method does not require access to LLM internals; instead, it leverages carefully designed prompt engineering to model the causal effect between log messages and their templates. To further refine the causal estimation, we incorporate counterfactual rewriting and retrieval-augmented generation (RAG) to enable accurate and unbiased inference over mediators. Empirical evaluation on multiple public log benchmark datasets demonstrates that DebiasParser achieves state-of-the-art performance, significantly improving both log grouping and template extraction accuracy. These findings validate the effectiveness of causal debiasing in enhancing LLM-based log parsing.