Rethinking the Reliability of Post-hoc Calibration Methods Under Subpopulation Shift
摘要
In recent decades, many researchers have directed their focus towards confidence calibration of deep neural networks because trustworthiness is equally important as accuracy, especially in high-stake scenarios within real-world applications. Unlike methods that involve modifying the training process, post-hoc calibration methods establish excellent performance by calibrating the model on the calibration set without fine-tuning the trained classifiers. Consequently, they are widely used in real-world applications. However, most post-hoc calibration methods assume that the distribution of calibration data is consistent with that of the test data and previous researches always conduct experiments on a fixed calibration set, while the situation in real-world applications is often different. Furthermore, the influence of the factors on calibration set for post-hoc calibration remains unclear. Therefore, a systematic investigation targeting the guidance for the deployment of different post-hoc calibration paradigms is required. In this study, we conduct experiments to evaluate the performance of post-hoc calibration paradigms under different settings on the calibration set and the effectiveness of different augmentation strategies under subpopulation shift.