Large language models (LLMs), as a significant direction in the development of artificial intelligence in recent years, are becoming increasingly popular. These models demonstrate significant potential in applications such as personalized services, simulated dialogues, role-playing, and specific compliance requirements. However, with the continuous expansion of LLMs’ practical applications, selective knowledge management of the models—particularly the ability to make the models “forget” certain domain-specific or topic-specific knowledge without impacting their overall performance or knowledge in other areas—has emerged as a critical issue that demands resolution. We proposes an innovative knowledge domain forgetting method designed to enable models to selectively forget specified knowledge without the need for fine-tuning. The method identifies and removes steering vectors associated with specific knowledge domains, thereby achieving effective forgetting of specific knowledge. The proposed approach has been evaluated in several popular open-source LLMs. Experimental results show that this method achieves good knowledge forgetting effects across diverse scenarios and exhibits notable practical value.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Removing Regional Steering Vectors to Achieve Knowledge Domain Forgetting in Large Language Models

  • Wei Wu,
  • Chen Wang,
  • Qiuhao Xu,
  • Wei Kong

摘要

Large language models (LLMs), as a significant direction in the development of artificial intelligence in recent years, are becoming increasingly popular. These models demonstrate significant potential in applications such as personalized services, simulated dialogues, role-playing, and specific compliance requirements. However, with the continuous expansion of LLMs’ practical applications, selective knowledge management of the models—particularly the ability to make the models “forget” certain domain-specific or topic-specific knowledge without impacting their overall performance or knowledge in other areas—has emerged as a critical issue that demands resolution. We proposes an innovative knowledge domain forgetting method designed to enable models to selectively forget specified knowledge without the need for fine-tuning. The method identifies and removes steering vectors associated with specific knowledge domains, thereby achieving effective forgetting of specific knowledge. The proposed approach has been evaluated in several popular open-source LLMs. Experimental results show that this method achieves good knowledge forgetting effects across diverse scenarios and exhibits notable practical value.