Artificial intelligence systems, particularly those utilizing deep neural networks, are increasingly integrated into critical sectors like healthcare, finance, and security. These systems’ reliance underscores the importance of explainable AI, commonly known as xAI, which aims to explain the decision-making processes of AI models. However, xAI’s integrity is challenged by adversarial attacks that manipulate explanations without affecting the model’s output. This paper introduces a novel method to generate such attacks, highlighting their potential to mislead operators and compromise system trustworthiness. By detailing the susceptibility of xAI to these explanation attacks, the study emphasizes the need for robustness in xAI methods. It proposes that ensuring the fidelity of explanations under adversarial conditions is crucial for maintaining the transparency and reliability of AI systems. Through this work, the authors aim to advance understanding of security vulnerabilities in xAI and contribute to the development of AI systems that are both transparent and resilient against adversarial threats. This investigation into the phenomenon of explanation attacks opens a new possibility for enhancing the security protocols surrounding AI systems in sensitive applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Proposition of a Novel Type of Attacks Targetting Explainable AI Algorithms in Cybersecurity

  • Sebastian Szelest,
  • Marek Pawlicki,
  • Aleksandra Pawlicka,
  • Rafał Kozik,
  • Michał Choraś

摘要

Artificial intelligence systems, particularly those utilizing deep neural networks, are increasingly integrated into critical sectors like healthcare, finance, and security. These systems’ reliance underscores the importance of explainable AI, commonly known as xAI, which aims to explain the decision-making processes of AI models. However, xAI’s integrity is challenged by adversarial attacks that manipulate explanations without affecting the model’s output. This paper introduces a novel method to generate such attacks, highlighting their potential to mislead operators and compromise system trustworthiness. By detailing the susceptibility of xAI to these explanation attacks, the study emphasizes the need for robustness in xAI methods. It proposes that ensuring the fidelity of explanations under adversarial conditions is crucial for maintaining the transparency and reliability of AI systems. Through this work, the authors aim to advance understanding of security vulnerabilities in xAI and contribute to the development of AI systems that are both transparent and resilient against adversarial threats. This investigation into the phenomenon of explanation attacks opens a new possibility for enhancing the security protocols surrounding AI systems in sensitive applications.