Quantifying the efficacy of scenario based jailbreaks on Llama using HarmBench and LLM-as-a-judge
摘要
While alignment techniques like RLHF and DPO protect large language models from direct malicious prompts, these systems remain susceptible to nuanced adversarial environments. In this study, these vulnerabilities are examined by benchmarking Llama-3.1-8B-Instruct against a baseline model, Gemma-2-9B-IT, and evaluating the models’ resistance to the contextual jailbreak attack type, in which an untrustworthy context is placed before a benign prompt. With the exhaustive HarmBench contextual dataset (