<p>We propose a method based on knowledge distillation to deal with adversarial attacks in chatbots, where malicious users deliberately input toxic dialogues to make the bot imitate their speech. The approach involves establishing a good teacher who only learns clean dialogues to increase the probability of clean dialogues being output by students and a nasty teacher who only knows toxic dialogues to reduce the probability of toxic dialogues. We find that the temperature hyperparameter for the good and nasty teachers should be set to 20 and 10, respectively, and the model with knowledge distillation can learn more knowledge than the model without distillation. The results show that the proposed bi-long short-term memory (LSTM) + dual teachers of knowledge distillation (DTKD) model helps improve dialogue ability. The study provides directions and suggestions for future research, including finding suitable datasets and extending the design of knowledge distillation architecture and toxicity detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Avoiding toxicity and prejudice in chatbots with knowledge distillation

  • Hei-Chia Wang,
  • Cendra Devayana Putra,
  • Hsin-Tzu Weng

摘要

We propose a method based on knowledge distillation to deal with adversarial attacks in chatbots, where malicious users deliberately input toxic dialogues to make the bot imitate their speech. The approach involves establishing a good teacher who only learns clean dialogues to increase the probability of clean dialogues being output by students and a nasty teacher who only knows toxic dialogues to reduce the probability of toxic dialogues. We find that the temperature hyperparameter for the good and nasty teachers should be set to 20 and 10, respectively, and the model with knowledge distillation can learn more knowledge than the model without distillation. The results show that the proposed bi-long short-term memory (LSTM) + dual teachers of knowledge distillation (DTKD) model helps improve dialogue ability. The study provides directions and suggestions for future research, including finding suitable datasets and extending the design of knowledge distillation architecture and toxicity detection.