<p>Convolutional Neural Networks (CNNs) have achieved tremendous success in image classification tasks. However, CNNs are vulnerable to adversarial attacks, such as applying imperceptible perturbations on the legitimate images. To address the security threats posed by these adversarial attacks, many defense techniques have been proposed. Adversarial training has been shown to be effective in enhancing CNNs robustness against adversarial samples. However, the trade-off between robustness and classification accuracy in adversarial training cannot be overlooked. In this paper, we propose a novel approach to adversarial training that simultaneously trains the model using both real images and minimally perturbed borderline adversaries. These borderline adversaries were generated during the training process, using the shortest successful perturbations for each individual training sample at specific training states. Instead of training with a fixed <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7477_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="10" /> </InlineMediaObject> <EquationSource Format="TEX">\(\epsilon\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ϵ</mi> </math></EquationSource> </InlineEquation> that applies uniform perturbations to all training samples, this shortest successful perturbation is adaptive to the network’s training state and is automatically determined. The rationale behind this approach is that the decision boundary will be less distorted by these additional adversaries, which helps maintain the classification accuracy while improving adversarial robustness. Preliminary experiments conducted on Cholec80 dataset for surgical tool recognition showed that this method achieved 2–7% improvement in both adversarial robustness and accuracy compared to other adversarial training methods, while also reducing overlap in the classification regions after adversarial training.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adversarial training with borderline samples

  • Ning Ding,
  • Knut Möller

摘要

Convolutional Neural Networks (CNNs) have achieved tremendous success in image classification tasks. However, CNNs are vulnerable to adversarial attacks, such as applying imperceptible perturbations on the legitimate images. To address the security threats posed by these adversarial attacks, many defense techniques have been proposed. Adversarial training has been shown to be effective in enhancing CNNs robustness against adversarial samples. However, the trade-off between robustness and classification accuracy in adversarial training cannot be overlooked. In this paper, we propose a novel approach to adversarial training that simultaneously trains the model using both real images and minimally perturbed borderline adversaries. These borderline adversaries were generated during the training process, using the shortest successful perturbations for each individual training sample at specific training states. Instead of training with a fixed \(\epsilon\) ϵ that applies uniform perturbations to all training samples, this shortest successful perturbation is adaptive to the network’s training state and is automatically determined. The rationale behind this approach is that the decision boundary will be less distorted by these additional adversaries, which helps maintain the classification accuracy while improving adversarial robustness. Preliminary experiments conducted on Cholec80 dataset for surgical tool recognition showed that this method achieved 2–7% improvement in both adversarial robustness and accuracy compared to other adversarial training methods, while also reducing overlap in the classification regions after adversarial training.