错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Diagnostics-Guided Explanation Generation

  • Pepa Atanasova

摘要

Explanations help reveal the reasoning behind a model’s predictions, building user trust and aiding in identifying vulnerabilities. European law even emphasizes the right to obtain an explanation for automated decisions (Regulation 2016). This is a summary of the Atanasova et al. (2021) work, which explores how the properties of sentence-level explanations for machine learning (ML) models trained for complex reasoning tasks can be automatically enhanced. When human explanation annotations are absent, a common approach is to train models that select regions from the input, maximizing proximity to the original task performance, known as the Faithfulness property (Lei et al. 2016; Yu et al. 2019). This chapter further explores other diagnostic properties, including Data Consistency and Confidence Indication. The contributions of the chapter include presenting the first method to learn diagnostic properties in an unsupervised manner, optimizing for them to enhance the quality of generated explanations. A joint model is implemented for task prediction and explanation generation, with each diagnostic property serving as an additional training objective. Experiments on three reasoning tasks show that optimizing for diagnostic properties not only improves those properties but also leads to explanations with higher human agreement and enhanced task performance. The chapter also highlights the importance of human rationales in training models for accurate predictions with meaningful explanations.