<p>The increasing prevalence of online user generated content has raised serious concerns about toxic language, which reinforces societal biases and causes psychological harm. This study introduces a novel approach that combines multi agent debate and reinforcement learning to improve the detoxification of large language models. The proposed framework enhances model robustness by enabling iterative refinement through agent interaction, effectively reducing toxicity, especially in contexts involving hate speech. In addition, reinforcement learning is applied to fine tune sequence to sequence models for detoxifying dialogue summaries. Using models such as FLAN T5, BART, and GODEL, we evaluate the approach on RealToxicity Prompts and ParaDetox datasets. The results show consistent reductions in toxicity scores while maintaining content fidelity and coherence. These findings demonstrate the effectiveness of multi agent collaboration and learning based adaptation in mitigating toxic language and improving safety in real world applications such as content moderation and summarization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detoxifying language model outputs: combining multi-agent debates and reinforcement learning for improved summarization

  • G. Bharathi Mohan,
  • M. Gayathri,
  • R. Prasanna Kumar

摘要

The increasing prevalence of online user generated content has raised serious concerns about toxic language, which reinforces societal biases and causes psychological harm. This study introduces a novel approach that combines multi agent debate and reinforcement learning to improve the detoxification of large language models. The proposed framework enhances model robustness by enabling iterative refinement through agent interaction, effectively reducing toxicity, especially in contexts involving hate speech. In addition, reinforcement learning is applied to fine tune sequence to sequence models for detoxifying dialogue summaries. Using models such as FLAN T5, BART, and GODEL, we evaluate the approach on RealToxicity Prompts and ParaDetox datasets. The results show consistent reductions in toxicity scores while maintaining content fidelity and coherence. These findings demonstrate the effectiveness of multi agent collaboration and learning based adaptation in mitigating toxic language and improving safety in real world applications such as content moderation and summarization.