Development of a Red-Teaming Dataset to Defend Large Language Models Against Attacks
摘要
Abstract
State-of-the-art (SOTA) large language models (LLMs) are huge systems with complex internal mechanisms that implement black-box response generation. Even though aligned LLMs have built-in defense mechanisms against attacks, recent studies demonstrate the vulnerability of LLMs to jailbreak attacks. In this study, we aim to extend the existing red-teaming datasets obtained from jailbreak attacks to address these LLM vulnerabilities in the future. In addition, we carry out some experiments with SOTA LLMs on our dataset to demonstrate the existing weaknesses in these models.