In the field of cybersecurity, the emergence of Large Language Models (LLMs) has opened up a new domain of potential risks. Although these models show an impressive level of capability across a range of applications, they are not free from the possibility of creating content that could negatively impact both social and digital safety. This paper examines the implications of In-Context Learning based command attacks, a burgeoning threat to the security and ethical integrity of LLMs. We introduce the In-Context Tense Attack (ITA) framework, a novel approach that employs harmful examples to undermine the integrity of LLMs. Our theoretical analysis elucidates how a constrained set of context examples can significantly influence the security mechanisms of LLMs. Through rigorous experimentation, we have substantiated the potency of ITA in elevating the success rate of jailbreaking prompts. On the PKU-Alignment/SafeRLHF dataset, ITA achieved a remarkable 92.99% increase in Accuracy, a 73.36% improvement in Rouge-L, and a 27.01% enhancement in Bleu-4 scores. Similarly, on the NVIDIA/Aegis-Safety dataset, ITA demonstrated a 72.03% increase in Accuracy, an 80.87% rise in Rouge-L, and a 40.24% boost in Bleu-4 scores. These results underscore the effectiveness of ITA in manipulating LLMs to generate harmful outputs, thereby highlighting the necessity for more robust security measures.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bypassing LLM Safeguards: The In-Context Tense Attack Approach

  • Shaohuang Wang,
  • Ruijing Geng,
  • Shuai Lei,
  • Yanfei Lv,
  • Huaping Zhang

摘要

In the field of cybersecurity, the emergence of Large Language Models (LLMs) has opened up a new domain of potential risks. Although these models show an impressive level of capability across a range of applications, they are not free from the possibility of creating content that could negatively impact both social and digital safety. This paper examines the implications of In-Context Learning based command attacks, a burgeoning threat to the security and ethical integrity of LLMs. We introduce the In-Context Tense Attack (ITA) framework, a novel approach that employs harmful examples to undermine the integrity of LLMs. Our theoretical analysis elucidates how a constrained set of context examples can significantly influence the security mechanisms of LLMs. Through rigorous experimentation, we have substantiated the potency of ITA in elevating the success rate of jailbreaking prompts. On the PKU-Alignment/SafeRLHF dataset, ITA achieved a remarkable 92.99% increase in Accuracy, a 73.36% improvement in Rouge-L, and a 27.01% enhancement in Bleu-4 scores. Similarly, on the NVIDIA/Aegis-Safety dataset, ITA demonstrated a 72.03% increase in Accuracy, an 80.87% rise in Rouge-L, and a 40.24% boost in Bleu-4 scores. These results underscore the effectiveness of ITA in manipulating LLMs to generate harmful outputs, thereby highlighting the necessity for more robust security measures.