Once a classification model has been implemented, adverse distribution shifts can surface. As true labels are costly to obtain or may only be known after a certain lag, cautious monitoring should be applied to prevent unnoticed and harmful deterioration of the model’s performance. In this paper, we consider distribution shift settings where the presence of shifts is uncertain and the target dataset contains only unlabeled examples. We suggest two methods for practitioners to assess the model’s performance in that context. Those techniques aim to exploit patterns which result in incorrect predictions by the model or unreliable uncertainty assessment. The Calibration Error Grid (CE Grid) exploits regional disparities in model calibration quality to adjust the model’s confidence on unlabeled data. The Error Classifier (EC) is a supervised learning approach which leverages representations from the data coupled with the model’s confidence in order to predict the probability of error. To prove their efficacy, these techniques are empirically tested over different scenarios of shifts on various use cases. These two techniques display promising results when compared against four crucial baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Harnessing Error Patterns to Estimate Out-of-Distribution Performance

  • Thomas Bonnier,
  • Benjamin Bosch

摘要

Once a classification model has been implemented, adverse distribution shifts can surface. As true labels are costly to obtain or may only be known after a certain lag, cautious monitoring should be applied to prevent unnoticed and harmful deterioration of the model’s performance. In this paper, we consider distribution shift settings where the presence of shifts is uncertain and the target dataset contains only unlabeled examples. We suggest two methods for practitioners to assess the model’s performance in that context. Those techniques aim to exploit patterns which result in incorrect predictions by the model or unreliable uncertainty assessment. The Calibration Error Grid (CE Grid) exploits regional disparities in model calibration quality to adjust the model’s confidence on unlabeled data. The Error Classifier (EC) is a supervised learning approach which leverages representations from the data coupled with the model’s confidence in order to predict the probability of error. To prove their efficacy, these techniques are empirically tested over different scenarios of shifts on various use cases. These two techniques display promising results when compared against four crucial baselines.