Subject Information Extraction for Novelty Detection with Domain Shifts
摘要
Unsupervised novelty detection (UND) is vital for applications such as medical diagnosis and cybersecurity. A prevalent assumption in current UND methods is that the normal data for training and testing are drawn from the same domain. This assumption is challenged in practice by domain shift, where a disparity exists between the training and testing domains (e.g., due to different data acquisition pipelines or sites). Domain shift often leads to the incorrect classification of normal data as novel by existing methods. A typical situation is that the normal testing data and normal training data describe the same underlying subject semantics, yet they differ in background/domain conditions. To address this problem, we introduce a novel method that separates subject information from background variation, encapsulating the domain information to enhance detection performance under domain shifts. The proposed method minimizes mutual information between the representations of the subject and background while modeling the background variation using a deep Gaussian mixture model; novelty detection is conducted solely on the subject representations, thus reducing sensitivity to domain variations. Experiments on Multi-background MNIST and Kurcuma demonstrate that our model generalizes effectively to unseen domains and outperforms strong baselines under domain shifts.