Safety, security, and compliance are essential requirements when aligning large language models (LLMs). To reduce harm and misuse, efforts have been made to align these LLMs to human values using advanced training techniques such as Reinforcement Learning from Human Feedback (RLHF). This chapter presents attacks and defenses developed for aligned LLMs. Efforts made into circumventing the safety guardrails of aligned LLMs are often called jailbreaks. In this context, attacks aim to find vulnerability of aligned LLMs to adversarial jailbreak attempts aiming at subverting the embedded safety guardrails, while defenses aim to either detect malicious prompts or enhance the refusal capability to attacks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attacks and Defenses on Aligned Large Language Models

  • Pin-Yu Chen,
  • Sijia Liu

摘要

Safety, security, and compliance are essential requirements when aligning large language models (LLMs). To reduce harm and misuse, efforts have been made to align these LLMs to human values using advanced training techniques such as Reinforcement Learning from Human Feedback (RLHF). This chapter presents attacks and defenses developed for aligned LLMs. Efforts made into circumventing the safety guardrails of aligned LLMs are often called jailbreaks. In this context, attacks aim to find vulnerability of aligned LLMs to adversarial jailbreak attempts aiming at subverting the embedded safety guardrails, while defenses aim to either detect malicious prompts or enhance the refusal capability to attacks.