Abstract <p>State-of-the-art (SOTA) large language models (LLMs) are huge systems with complex internal mechanisms that implement black-box response generation. Even though aligned LLMs have built-in defense mechanisms against attacks, recent studies demonstrate the vulnerability of LLMs to jailbreak attacks. In this study, we aim to extend the existing red-teaming datasets obtained from jailbreak attacks to address these LLM vulnerabilities in the future. In addition, we carry out some experiments with SOTA LLMs on our dataset to demonstrate the existing weaknesses in these models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a Red-Teaming Dataset to Defend Large Language Models Against Attacks

  • I. S. Alekseevskaia,
  • K. V. Arkhipenko,
  • D. Yu. Turdakov

摘要

Abstract

State-of-the-art (SOTA) large language models (LLMs) are huge systems with complex internal mechanisms that implement black-box response generation. Even though aligned LLMs have built-in defense mechanisms against attacks, recent studies demonstrate the vulnerability of LLMs to jailbreak attacks. In this study, we aim to extend the existing red-teaming datasets obtained from jailbreak attacks to address these LLM vulnerabilities in the future. In addition, we carry out some experiments with SOTA LLMs on our dataset to demonstrate the existing weaknesses in these models.