While previous work in explainable artificial intelligence has indicated potential relevance of conversational (natural language dialogue-based) explainability, we lack a comprehensive analysis of dialogue capabilities that may be desired from conversationally explainable (CXAI) systems, as well as methods and analytical tools for assessing and comparing conversational capabilities of such systems. To bridge this gap, we develop a taxonomy of explanatory dialogue capabilities, encompassing 19 items pertaining to question answering and information delivery, context management, and grounding and meta-communication. The benchmark complements end-user testing by enabling limitations and areas of improvement to be detected at an early stage, with a relatively small effort. We apply the proposed benchmark to three CXAI systems: TalkToModel [27], Glass-Box [28], and BKOS [3]. The results reveal gaps in supported capabilities across all studied systems, and a substantial amount of variation across the systems, indicating ample room for future work in this area. We also present methodological recommendations for facilitating continued assessment and progress by means of automated dialogue testing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing Conversational Capabilities of Explanatory AI Interfaces

  • Alexander Berman,
  • Staffan Larsson

摘要

While previous work in explainable artificial intelligence has indicated potential relevance of conversational (natural language dialogue-based) explainability, we lack a comprehensive analysis of dialogue capabilities that may be desired from conversationally explainable (CXAI) systems, as well as methods and analytical tools for assessing and comparing conversational capabilities of such systems. To bridge this gap, we develop a taxonomy of explanatory dialogue capabilities, encompassing 19 items pertaining to question answering and information delivery, context management, and grounding and meta-communication. The benchmark complements end-user testing by enabling limitations and areas of improvement to be detected at an early stage, with a relatively small effort. We apply the proposed benchmark to three CXAI systems: TalkToModel [27], Glass-Box [28], and BKOS [3]. The results reveal gaps in supported capabilities across all studied systems, and a substantial amount of variation across the systems, indicating ample room for future work in this area. We also present methodological recommendations for facilitating continued assessment and progress by means of automated dialogue testing.