Removing Regional Steering Vectors to Achieve Knowledge Domain Forgetting in Large Language Models
摘要
Large language models (LLMs), as a significant direction in the development of artificial intelligence in recent years, are becoming increasingly popular. These models demonstrate significant potential in applications such as personalized services, simulated dialogues, role-playing, and specific compliance requirements. However, with the continuous expansion of LLMs’ practical applications, selective knowledge management of the models—particularly the ability to make the models “forget” certain domain-specific or topic-specific knowledge without impacting their overall performance or knowledge in other areas—has emerged as a critical issue that demands resolution. We proposes an innovative knowledge domain forgetting method designed to enable models to selectively forget specified knowledge without the need for fine-tuning. The method identifies and removes steering vectors associated with specific knowledge domains, thereby achieving effective forgetting of specific knowledge. The proposed approach has been evaluated in several popular open-source LLMs. Experimental results show that this method achieves good knowledge forgetting effects across diverse scenarios and exhibits notable practical value.