Adversarials: Anti-AI Countermeasures
摘要
A battlefield is an adversarial space, and opponents will prepare countermeasures against autonomous weapon systems (AWS). Unlike countermeasures against older systems, AI have unique vulnerabilities that can be exploited in novel ways, which must be understood by any AWS-user to prevent opponents from compromising the system’s reliability, force and civilian safety, and mission success. This chapter features an in-depth exploration of adversarials—anti-AI countermeasures. The chapter first distinguishes adversarials from ‘traditional’ countermeasures, such as integrity attacks and jamming, by identifying the property unique to modern AI that adversarials exploit: data reliance. From this, three main types of adversarials are analysed. Adversarial inputs manipulate data the AWS uses to make decisions (e.g. through visual patches or perturbed ratio waves) in order to induce a particular behaviour, e.g. friendly-firing or shooting civilians. Poisoning compromises the AWS in its design phase, by manipulating the training data or model prior to development. Finally, backdoors are found to be a combination of the two above, and allow very sophisticated manipulation. From this analysis, it is found that much of the responsibility to collect relevant information and take the necessary precautions to mitigate risk from adversarials lies with the end-user, since it requires the construction of a threat model unique to each adversary. To assist in this endeavour, the chapter concludes by providing a guide to conduct security analysis based on the opponent’s goals, incentives, knowledge and power, which should inform commanders in the field about the risk of adversarials and possible counterplay.