A fundamental issue with machine learning is that the training data must contain sufficient examples of every data pattern of interest for the essentially statistical techniques of machine learning to derive a model that accurately covers all data patterns. For example, in real-world situations, samples from minority populations may not be adequately represented in available training datasets, resulting in models that are biased in favour of majority populations. We present an approach to partially address this problem by using domain knowledge to construct rules to correctly classify, specifically, those cases representing rare patterns in training data that the machine learning model has incorrectly classified. We use Ripple-Down Rules (RDR) to add rules, as it is a proven approach to knowledge acquisition and incremental maintenance that enables rules to be very easily and quickly added to a knowledge base on a case-by-case basis. This paper extends previous work using a human expert by using both artificial and real-world data and a simulated expert to add rules in multiple controlled experiments. Results show that domain knowledge can be successfully used in our approach to mitigate performance issues arising from a limited representation of a subgroup in the available training dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Responsible Decisions with Limited Training Data Using Human-in-the-Loop

  • Ashesh Mahidadia,
  • Michael Bain,
  • Hendra Suryanto,
  • Byeong Kang,
  • Charles Guan,
  • Paul Compton

摘要

A fundamental issue with machine learning is that the training data must contain sufficient examples of every data pattern of interest for the essentially statistical techniques of machine learning to derive a model that accurately covers all data patterns. For example, in real-world situations, samples from minority populations may not be adequately represented in available training datasets, resulting in models that are biased in favour of majority populations. We present an approach to partially address this problem by using domain knowledge to construct rules to correctly classify, specifically, those cases representing rare patterns in training data that the machine learning model has incorrectly classified. We use Ripple-Down Rules (RDR) to add rules, as it is a proven approach to knowledge acquisition and incremental maintenance that enables rules to be very easily and quickly added to a knowledge base on a case-by-case basis. This paper extends previous work using a human expert by using both artificial and real-world data and a simulated expert to add rules in multiple controlled experiments. Results show that domain knowledge can be successfully used in our approach to mitigate performance issues arising from a limited representation of a subgroup in the available training dataset.