Few-Shot Unlearning for Large Language Models Through Pretrained Soft Prompts
摘要
We propose FLARE, a novel few-shot unlearning framework for LLMs. FLARE conceptualizes the unlearning tasks as relearning tasks, focusing on learning the new labels for the unwanted instances. FLARE integrates task-specific and instance-specific information into a shared source prompt called TI prompt and adapts TI prompt to downstream unlearning tasks. We conduct the experiments across different few-shot unlearning tasks, and the results demonstrate that FLARE significantly outperforms state-of-the-art baselines despite fine-tuning only 1% to 2% of the LLM’s parameters.