错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Introduction

  • Aaron M. Roth,
  • Dinesh Manocha,
  • Ram D. Sriram,
  • Elham Tabassi

摘要

We survey the state of the art in Explainable and Interpretable Reinforcement Learning as relevant for Robotics. We propose a classification system to help facilitate discussion of and research into this interdisciplinary field. We consider classification techniques used in past surveys and papers and attempt to unify terminology, describing 12 attributes that can be used to classify explainable/interpretable techniques: (i) Model-Agnostic or Model-Specific (MAMS), (ii) Self-Explainable or Post-Hoc (SEPH), (iii) Scope, (iv) When-Produced, (v) Format, (vi) Knowledge Limits, (vii) Explanation Accuracy, (viii) Audience, (ix) Predictability, (x) Legibility, (xi) Readability, and (xii) Reactivity. We organize our discussion of methods into 42 categories and subcategories, each classified according to some of the attributes. The contributions of this survey include; (1) a review of how related terminology is used in the field and past proposed classification systems, (2) a novel classification system to classify prior work into 12 important attributes, identifying consensus where such exists and identifying the best existing terminology otherwise, (3) a broad survey of the state of the art in XRL for robotics, divided into categories and subcategories, with descriptions of some methods, and subcategories or categories described according to the attributes noted, and (4) Identification and proposals of areas for future research, including specific suggestions for research projects.