Bypassing LLM Safeguards: The In-Context Tense Attack Approach
摘要
In the field of cybersecurity, the emergence of Large Language Models (LLMs) has opened up a new domain of potential risks. Although these models show an impressive level of capability across a range of applications, they are not free from the possibility of creating content that could negatively impact both social and digital safety. This paper examines the implications of In-Context Learning based command attacks, a burgeoning threat to the security and ethical integrity of LLMs. We introduce the In-Context Tense Attack (ITA) framework, a novel approach that employs harmful examples to undermine the integrity of LLMs. Our theoretical analysis elucidates how a constrained set of context examples can significantly influence the security mechanisms of LLMs. Through rigorous experimentation, we have substantiated the potency of ITA in elevating the success rate of jailbreaking prompts. On the PKU-Alignment/SafeRLHF dataset, ITA achieved a remarkable 92.99% increase in Accuracy, a 73.36% improvement in Rouge-L, and a 27.01% enhancement in Bleu-4 scores. Similarly, on the NVIDIA/Aegis-Safety dataset, ITA demonstrated a 72.03% increase in Accuracy, an 80.87% rise in Rouge-L, and a 40.24% boost in Bleu-4 scores. These results underscore the effectiveness of ITA in manipulating LLMs to generate harmful outputs, thereby highlighting the necessity for more robust security measures.