Talk the Talk, Debate the Bias: LLM Alignment via Role-Play Rumble
摘要
ABias in LLMs can undermine user experiences and societal outcomes, yet existing mitigation methods depend heavily on human feedback, lack generalizability, and often produce overconfident or erratic responses. To address these challenges, we propose RLDF (Reinforcement Learning from Multi-role Debates as Feedback), a framework that leverages LLM-driven debates to generate both high- and low-bias examples for reward-model training. In self-reflection mode, the model debates itself; in teacher–student mode, a more advanced LLM guides the discussion. Experiments across multiple LLMs demonstrate that RLDF substantially reduces bias.