We provide an overview of recent approaches to extract simpler abstractions of complex neural networks using Angluin’s exact learning framework with queries and counterexamples. These simpler models approximate parts of the original model by focusing on a relevant collection of inputs/outputs. The aim of constructing such abstractions is to obtain high level information about machine learning models, which can be useful to detect harmful biases and other issues. We focus on concept classes applied for actively learning from machine learning models within Angluin’s framework, namely, automata and Horn logic. We also discuss approaches from the literature to extract decision trees from neural networks. We highlight the benefits and drawbacks of these approaches. Finally, we discuss promising possible next steps and applications of these approaches for extracting high level information from machine learning models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Actively Learning from Machine Learning Models with Queries and Counterexamples

  • Ana Ozaki

摘要

We provide an overview of recent approaches to extract simpler abstractions of complex neural networks using Angluin’s exact learning framework with queries and counterexamples. These simpler models approximate parts of the original model by focusing on a relevant collection of inputs/outputs. The aim of constructing such abstractions is to obtain high level information about machine learning models, which can be useful to detect harmful biases and other issues. We focus on concept classes applied for actively learning from machine learning models within Angluin’s framework, namely, automata and Horn logic. We also discuss approaches from the literature to extract decision trees from neural networks. We highlight the benefits and drawbacks of these approaches. Finally, we discuss promising possible next steps and applications of these approaches for extracting high level information from machine learning models.