Explainable Reinforcement Learning (XRL) has emerged as a critical subfield at the intersection of reinforcement learning (RL) and Explainable Artificial Intelligence (XAI), aiming to render the decision-making processes of learning agents interpretable, transparent, and accessible to human users. This paper introduces a comprehensive evaluation framework, the XRL H-F-E Metrics, to assess the human-friendliness of explanations generated by XRL systems. Drawing from interdisciplinary literature in computer science, cognitive psychology, philosophy of science, and human-computer interaction, the framework is structured across four dimensions: foundational principles (e.g., correctness, robustness, bias mitigation), cognitively aligned explanation types (e.g., “why”, “why not”, counterfactuals), characteristics of “good” explanations (e.g., contrastiveness, selectivity, causality), and human-friendly presentation attributes (e.g., comprehensibility, interactivity, personalization). This checklist provides both a theoretical model and a practical tool for fostering transparency and trust in RL applications, while also identifying key directions for future research, including quantitative metrics, adaptive explanations, and emotionally responsive interfaces.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human-Friendly Explanations Checklist for Reinforcement Learning: XRL H-F-E Checklist

  • Daniel Adrián Contreras Olivas,
  • Lourdes Martinez-Villaseñor

摘要

Explainable Reinforcement Learning (XRL) has emerged as a critical subfield at the intersection of reinforcement learning (RL) and Explainable Artificial Intelligence (XAI), aiming to render the decision-making processes of learning agents interpretable, transparent, and accessible to human users. This paper introduces a comprehensive evaluation framework, the XRL H-F-E Metrics, to assess the human-friendliness of explanations generated by XRL systems. Drawing from interdisciplinary literature in computer science, cognitive psychology, philosophy of science, and human-computer interaction, the framework is structured across four dimensions: foundational principles (e.g., correctness, robustness, bias mitigation), cognitively aligned explanation types (e.g., “why”, “why not”, counterfactuals), characteristics of “good” explanations (e.g., contrastiveness, selectivity, causality), and human-friendly presentation attributes (e.g., comprehensibility, interactivity, personalization). This checklist provides both a theoretical model and a practical tool for fostering transparency and trust in RL applications, while also identifying key directions for future research, including quantitative metrics, adaptive explanations, and emotionally responsive interfaces.