Understanding why language models (LMs) make specific decisions is crucial for their deployment in critical applications such as healthcare and legal systems, where transparency is essential. Social science research suggests that humans often explain decisions by contrasting them with alternative outcomes. While prior work has examined contrastive explanations for LMs, there has been limited investigation into their faithfulness and plausibility—key factors in evaluating explanation quality. This study addresses this gap by assessing contrastive explanations based on these criteria. We conduct extensive experiments comparing contrastive and non-contrastive explanations across various model sizes and tasks. In addition, we collect human rationales to evaluate their alignment with model-generated ones. Our findings reveal that contrastive explanations not only align better with human reasoning but also more accurately reflect the model’s decision-making. We also observe a positive correlation between model size and the effectiveness of contrastive explanations, although this relationship varies across tasks. These results underscore the potential of contrastive explanations to enhance the interpretability and reliability of LM outputs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Contrastive and Non-contrastive Explanations for Language Models

  • Yousra Chahinez Hadj Azzem,
  • Fouzi Harrag,
  • Ladjel Bellatreche

摘要

Understanding why language models (LMs) make specific decisions is crucial for their deployment in critical applications such as healthcare and legal systems, where transparency is essential. Social science research suggests that humans often explain decisions by contrasting them with alternative outcomes. While prior work has examined contrastive explanations for LMs, there has been limited investigation into their faithfulness and plausibility—key factors in evaluating explanation quality. This study addresses this gap by assessing contrastive explanations based on these criteria. We conduct extensive experiments comparing contrastive and non-contrastive explanations across various model sizes and tasks. In addition, we collect human rationales to evaluate their alignment with model-generated ones. Our findings reveal that contrastive explanations not only align better with human reasoning but also more accurately reflect the model’s decision-making. We also observe a positive correlation between model size and the effectiveness of contrastive explanations, although this relationship varies across tasks. These results underscore the potential of contrastive explanations to enhance the interpretability and reliability of LM outputs.