<p>Removing the influence of specific training samples from a trained classifier on demand, without retraining from scratch, is central to compliance with data-deletion requests. We study this problem under a trigger-based framing in which a visible pattern inserted into an input activates a forgetting behavior at inference time, so that the clean form of the input is classified normally while the triggered form yields a maximally uncertain response. We develop a three-stage training pipeline consisting of a Pretraining Stage on clean retain data, a Conditioning Stage that binds the trigger to a uniform-posterior response under a bounded Kullback-Leibler objective with joint batches, and an Unlearning Stage that applies the same uniform target to triggered forget samples while an Elastic Weight Consolidation penalty protects retention-critical parameters. We evaluate the method on CIFAR-10 and SVHN with a six-condition protocol that covers every combination of data split and trigger state and we report multi-seed results. On held-out test data the proposed model attains a conditional gap of 23.04 pp on CIFAR-10 and 71.54 pp on SVHN, while a baseline trained without trigger conditioning exhibits gaps of only 1.16 pp and 0.35 pp, confirming that the effect is a learned input-conditioned behavior rather than a property of the trigger pattern itself. Membership inference AUC is 0.502 on CIFAR-10 and 0.511 on SVHN, closer to chance than for the baseline. Selectivity in this framework is carried by the trigger state at inference rather than by a sample-level discrimination between forget and retain subsets, a structural property of any method whose input-side signal does not depend on sample identity. A residual clean-input capacity loss of roughly 17 pp on CIFAR-10 and 9 pp on SVHN remains the principal limitation, and the six-condition protocol is provided as a reusable tool for evaluating trigger-conditional unlearning methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Trigger-based conditional selective unlearning with elastic weight consolidation

  • Hyun Kwon,
  • Joo Bon Maeng,
  • Dae-Jin Kim

摘要

Removing the influence of specific training samples from a trained classifier on demand, without retraining from scratch, is central to compliance with data-deletion requests. We study this problem under a trigger-based framing in which a visible pattern inserted into an input activates a forgetting behavior at inference time, so that the clean form of the input is classified normally while the triggered form yields a maximally uncertain response. We develop a three-stage training pipeline consisting of a Pretraining Stage on clean retain data, a Conditioning Stage that binds the trigger to a uniform-posterior response under a bounded Kullback-Leibler objective with joint batches, and an Unlearning Stage that applies the same uniform target to triggered forget samples while an Elastic Weight Consolidation penalty protects retention-critical parameters. We evaluate the method on CIFAR-10 and SVHN with a six-condition protocol that covers every combination of data split and trigger state and we report multi-seed results. On held-out test data the proposed model attains a conditional gap of 23.04 pp on CIFAR-10 and 71.54 pp on SVHN, while a baseline trained without trigger conditioning exhibits gaps of only 1.16 pp and 0.35 pp, confirming that the effect is a learned input-conditioned behavior rather than a property of the trigger pattern itself. Membership inference AUC is 0.502 on CIFAR-10 and 0.511 on SVHN, closer to chance than for the baseline. Selectivity in this framework is carried by the trigger state at inference rather than by a sample-level discrimination between forget and retain subsets, a structural property of any method whose input-side signal does not depend on sample identity. A residual clean-input capacity loss of roughly 17 pp on CIFAR-10 and 9 pp on SVHN remains the principal limitation, and the six-condition protocol is provided as a reusable tool for evaluating trigger-conditional unlearning methods.