Ctta: a novel chain-of-thought transfer adversarial attacks framework for large language models
摘要
Recent studies have indicated that large language models (LLMs) remain susceptible to adversarial attacks, despite enhanced robustness through the chain-of-thought (CoT) capability. However, this capability also introduces the potential for more covert and effective adversarial attack methods. This paper proposes a CoT Transfer Adversarial attack framework (CTTA) for general LLMs. Initially, we utilize a pre-trained model based on the transformer architecture and fine-tune it on various tasks to serve as a surrogate model. Subsequently, different levels of adversarial attack algorithms are utilized, and the generated adversarial samples are used as transfer samples. A thought chain-based adversarial transfer attack framework is constructed using transfer samples and thought chain techniques. Finally, various indicators are utilized to assess the performance of the general LLMs in response to this attack. The results demonstrate that the attack framework surpasses current state-of-the-art research. Numerous experiments on LLMs with varying performance and parameter sizes have validated the effectiveness, stability, and generalizability of this attack. The model’s error response and the superiority of this attack are thoroughly examined using attention by gradient technology, confirming the security threats posed by LLMs when leveraging CoT capability. This has significant implications for enhancing the security and robustness of LLMs.