Abstract <p>Backdoor attacks pose a severe threat to the integrity of machine learning models, especially in real-world image classification tasks. In such attacks, adversaries embed malicious behaviors triggered by specific patterns in the training data, causing models to misclassify whenever the trigger is present. This paper introduces a novel, <i>model-agnostic</i> defense that systematically detects and removes backdoor-infected samples using a synergy of dimensionality reduction and unsupervised clustering. Unlike most existing methods that address <i>digitally</i> added triggers, our approach specifically targets <i>physically</i> embedded triggers (e.g., a bandage placed on a face), which closely resemble real-world occlusions and are therefore harder to detect. We first extract high-level features from a trusted, pre-trained model, reduce the feature dimensionality via Principal Component Analysis (PCA), and then fit Gaussian Mixture Models (GMMs) to cluster suspicious samples. By identifying and filtering out outlying clusters, we effectively isolate poisoned images without assuming knowledge of the trigger or requiring access to the victim model. Extensive experiments on face versus non-face classification demonstrate that our defense substantially reduces attack success rates while preserving high accuracy on clean data, offering a practical and robust solution against challenging backdoor scenarios.</p> Graphic Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Effective defense against physically embedded backdoor attacks via clustering-based filtering

  • Mohammed Kutbi

摘要

Abstract

Backdoor attacks pose a severe threat to the integrity of machine learning models, especially in real-world image classification tasks. In such attacks, adversaries embed malicious behaviors triggered by specific patterns in the training data, causing models to misclassify whenever the trigger is present. This paper introduces a novel, model-agnostic defense that systematically detects and removes backdoor-infected samples using a synergy of dimensionality reduction and unsupervised clustering. Unlike most existing methods that address digitally added triggers, our approach specifically targets physically embedded triggers (e.g., a bandage placed on a face), which closely resemble real-world occlusions and are therefore harder to detect. We first extract high-level features from a trusted, pre-trained model, reduce the feature dimensionality via Principal Component Analysis (PCA), and then fit Gaussian Mixture Models (GMMs) to cluster suspicious samples. By identifying and filtering out outlying clusters, we effectively isolate poisoned images without assuming knowledge of the trigger or requiring access to the victim model. Extensive experiments on face versus non-face classification demonstrate that our defense substantially reduces attack success rates while preserving high accuracy on clean data, offering a practical and robust solution against challenging backdoor scenarios.

Graphic Abstract