ABias in LLMs can undermine user experiences and societal outcomes, yet existing mitigation methods depend heavily on human feedback, lack generalizability, and often produce overconfident or erratic responses. To address these challenges, we propose RLDF (Reinforcement Learning from Multi-role Debates as Feedback), a framework that leverages LLM-driven debates to generate both high- and low-bias examples for reward-model training. In self-reflection mode, the model debates itself; in teacher–student mode, a more advanced LLM guides the discussion. Experiments across multiple LLMs demonstrate that RLDF substantially reduces bias.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Talk the Talk, Debate the Bias: LLM Alignment via Role-Play Rumble

  • Ruoxi Cheng,
  • Zhiqiang Wang,
  • Shaowei Yuan,
  • Yizhong Ding,
  • Rui Zhang

摘要

ABias in LLMs can undermine user experiences and societal outcomes, yet existing mitigation methods depend heavily on human feedback, lack generalizability, and often produce overconfident or erratic responses. To address these challenges, we propose RLDF (Reinforcement Learning from Multi-role Debates as Feedback), a framework that leverages LLM-driven debates to generate both high- and low-bias examples for reward-model training. In self-reflection mode, the model debates itself; in teacher–student mode, a more advanced LLM guides the discussion. Experiments across multiple LLMs demonstrate that RLDF substantially reduces bias.