A Survey of Image Captioning Adversarial Attacks
摘要
Image captioning, an essential research direction in the field of computer vision, finds extensive applications in real life. It refers to the process of converting images into textual descriptions, often encompassing information about objects, scenes, actions, and more within the image. This technology holds immense value in areas such as social media, search engines, and automated assistance. However, due to the vulnerability of deep neural networks, the application process is highly susceptible to adversarial examples, posing significant security risks and challenges. Adversarial attacks in the image domain have long been a popular research topic, and investigating them is crucial for enhancing security. Researchers in academia have conducted extensive research from various angles. Nevertheless, due to the multimodal nature involved, research on adversarial attacks specifically targeting image captioning remains scarce. Therefore, this review first introduces the basic concepts of image adversarial attacks. It then categorizes and summarizes existing adversarial attack methods for images based on their attack philosophies, introduces emerging attack algorithms targeting image captioning. Finally, the review summarizes the current research status, outlines the challenges faced by adversarial attacks in the field of image captioning and offers insights into future research directions.