Reinforcement Learning from Human Feedback (RLHF) is increasingly popular for aligning machine learning models’ outputs with human values, such as honesty, helpfulness, and harmlessness. Parameter Efficient Fine-Tuning (PEFT) has also shown promising outcomes, enabling small-sized models to produce results that are on par with huge models. In this paper, the power of RLHF and PEFT has been leveraged to fine-tune the FLAN-T5 base model on the DialogSum dataset for positive dialogue summarization tasks. It was demonstrated through quantitative and qualitative analysis that FLAN-T5 fine-tuned with PEFT and RLHF while maintaining the KL-Divergence metric to avoid hallucinations, outperforms a simple FLAN-T5 in producing positive and quality summaries. Meta’s variant of BERT—RoBERTa, as the reward model, Low-Rank Adaptation (LoRA) as the PEFT technique, and Proximal Policy Optimization (PPO) as the reinforcement algorithm were used. The mean and standard deviation of toxicity scores and the ROUGE scores were used to compare the models. The results suggest that RLHF combined with PEFT addresses challenges such as large computational requirements, catastrophic forgetting, unethical responses, and hallucinations in dialogue summarization tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging RLHF and PEFT to Fine-Tune FLAN-T5 for Generating Non-toxic/Positive Summaries

  • Riya Ahlawat,
  • Muskan Singh,
  • Samriddhi Tiwari,
  • Shweta Jindal

摘要

Reinforcement Learning from Human Feedback (RLHF) is increasingly popular for aligning machine learning models’ outputs with human values, such as honesty, helpfulness, and harmlessness. Parameter Efficient Fine-Tuning (PEFT) has also shown promising outcomes, enabling small-sized models to produce results that are on par with huge models. In this paper, the power of RLHF and PEFT has been leveraged to fine-tune the FLAN-T5 base model on the DialogSum dataset for positive dialogue summarization tasks. It was demonstrated through quantitative and qualitative analysis that FLAN-T5 fine-tuned with PEFT and RLHF while maintaining the KL-Divergence metric to avoid hallucinations, outperforms a simple FLAN-T5 in producing positive and quality summaries. Meta’s variant of BERT—RoBERTa, as the reward model, Low-Rank Adaptation (LoRA) as the PEFT technique, and Proximal Policy Optimization (PPO) as the reinforcement algorithm were used. The mean and standard deviation of toxicity scores and the ROUGE scores were used to compare the models. The results suggest that RLHF combined with PEFT addresses challenges such as large computational requirements, catastrophic forgetting, unethical responses, and hallucinations in dialogue summarization tasks.