This paper proposes an LLM guardrail framework that incorporates a Zero Trust architecture to validate and control the responses of Large Language Model (LLM) to unethical queries. The proposed framework applies guardrails to harmful inputs to avoid harmful responses and includes four verification steps through Policy Decision Point (PDP) and Policy Enforcement Point (PEP) structures. This structure aims to enhance the reliability and safety of LLM responses. We demonstrate this framework on a fixed model and verify its generality by applying it to various models. Consequently, this allows for evasive responses to a wide range of unethical prompts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM Guardrail Framework: A Novel Approach for Implementing Zero Trust Architecture

  • Bogeum Kim,
  • Hyejin Sim,
  • Jiwon Yun,
  • Jaehan Cho,
  • Howon Kim

摘要

This paper proposes an LLM guardrail framework that incorporates a Zero Trust architecture to validate and control the responses of Large Language Model (LLM) to unethical queries. The proposed framework applies guardrails to harmful inputs to avoid harmful responses and includes four verification steps through Policy Decision Point (PDP) and Policy Enforcement Point (PEP) structures. This structure aims to enhance the reliability and safety of LLM responses. We demonstrate this framework on a fixed model and verify its generality by applying it to various models. Consequently, this allows for evasive responses to a wide range of unethical prompts.