Classification System
摘要
This chapter details our Classification System and the rationale behind it. We begin by discussing prior surveys on Explainable Reinforcement Learning or Explainable Robotics, and any previous attempts at taxonomizing these subfields. Our own system involves twelve attributes which are grouped into three types of attributes. Hard Attributes are attributes that can generally be objectively defined. Hard Attributes include (i) MAMS (is the method model-agnostic or model-specific?) (ii) SEPH (does the method produce intrinsically self-explainable policy or is a process applied post-hoc?), (iii) Scope (global or local), (iv) When-Produced (is explanation produced before, during, or after training? If during, is it an intrinsic attribute of policy or a byproduct?) and (v) Format (i.e. text, image, rules, lists, numerical, etc). Soft Attributes are attributes that are imprecisely defined. Soft Attributes can be robot-specific or general. General Soft Attributes include (vi) Knowledge Limits (does system understand its own applicability and limits?), (vii) Explanation Accuracy (how accurate are explanations themselves), and (viii) Audience. Robot-specific Soft Attributes include (ix) Predictability, (x) Legibility (does robot behavior imply its goals), (xi) Readability (does robot behavior imply its next action), and (xii) Reactivity (does the robot react to an environment, plan deliberatively, or some combination).