Learning composite concepts, such as “red car”, from individual examples—like a white car representing the concept of “car” and a red strawberry representing the concept of “red”—is inherently challenging. This paper introduces a novel method called Composite Concept Extractor (CoCE), which leverages techniques from traditional backdoor attacks to learn these composite concepts in a zero-shot setting, requiring only examples of individual concepts. By repurposing the trigger-based model backdooring mechanism, we create a strategic distortion in the manifold of the target object (e.g., “car”) induced by example objects with the target property (e.g., “red”) from objects “red strawberry”, ensuring the distortion selectively affects the target objects with the target property. Contrastive learning is then employed to further refine this distortion and a method is formulated for detecting objects that are influenced by the distortion. Extensive experiments with an in-depth analysis across different datasets demonstrates the utility and applicability of our proposed approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Composite Concept Extraction Through Backdooring

  • Banibrata Ghosh,
  • Haripriya Harikumar,
  • Khoa D. Doan,
  • Svetha Venkatesh,
  • Santu Rana

摘要

Learning composite concepts, such as “red car”, from individual examples—like a white car representing the concept of “car” and a red strawberry representing the concept of “red”—is inherently challenging. This paper introduces a novel method called Composite Concept Extractor (CoCE), which leverages techniques from traditional backdoor attacks to learn these composite concepts in a zero-shot setting, requiring only examples of individual concepts. By repurposing the trigger-based model backdooring mechanism, we create a strategic distortion in the manifold of the target object (e.g., “car”) induced by example objects with the target property (e.g., “red”) from objects “red strawberry”, ensuring the distortion selectively affects the target objects with the target property. Contrastive learning is then employed to further refine this distortion and a method is formulated for detecting objects that are influenced by the distortion. Extensive experiments with an in-depth analysis across different datasets demonstrates the utility and applicability of our proposed approach.