Showing Many Labels in Multi-label Classification Models: An Empirical Study of Adversarial Examples
摘要
Deep Neural Networks (DNNs) have been widely applied in various fields. However, research indicates that DNNs are susceptible to adversarial examples in both multi-class and multi-label domains. To better understand existing multi-label adversarial attacks, this paper introduces a new attack type called “Showing Many Labels”. In this type of attack, the attacker’s goal is to maximize the number of labels in the model’s prediction results. Research in this area is significant as attackers could force AI systems to output additional labels, potentially causing serious consequences, such as incorrect decisions in autonomous driving systems. To investigate the ability of existing attack methods to increase the number of labels predicted by a model, we evaluate the attack success rates and the magnitude of generated perturbations for nine attack methods under the “Showing Many Labels” setting, using two target models and four datasets. The work in this paper demonstrates that existing attack methods can mislead models into identifying all labels from the dataset in adversarial examples, while remaining stealthy. This indicates that achieving “Showing Many Labels” in real-world scenarios is possible.