This paper describes the application of knowledge grounded generation for the personification of dialogue agents in Russian language. To implement own model for Russian language the authors propose two different architectures for personified systems: BERT&GPT and T5-multitask. The former consists of two independent components - a ranking module using an encoder-only BERT-style model and a generative module using a GPT model. The latter employs a unified multitask encoder-decoder model for both ranking and generation tasks. The research is based on the crowdsourced dataset Toloka Persona Chat Rus pre- processed to match the personification task. The authors used full golden knowledge set during training and hyperparameter tuning to avoid noise and improve quality of the knowledge-generation task. Manual evaluation of the proposed models’ metrics demonstrates a significant improvement in SSA scores for the Russian language, outperforming base models by 51%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid Approach to the Personification of Dialogue Agents

  • Pavel Posokhov,
  • Stepan Skrylnikov,
  • Olesia Makhnytkina,
  • Yuri Matveev

摘要

This paper describes the application of knowledge grounded generation for the personification of dialogue agents in Russian language. To implement own model for Russian language the authors propose two different architectures for personified systems: BERT&GPT and T5-multitask. The former consists of two independent components - a ranking module using an encoder-only BERT-style model and a generative module using a GPT model. The latter employs a unified multitask encoder-decoder model for both ranking and generation tasks. The research is based on the crowdsourced dataset Toloka Persona Chat Rus pre- processed to match the personification task. The authors used full golden knowledge set during training and hyperparameter tuning to avoid noise and improve quality of the knowledge-generation task. Manual evaluation of the proposed models’ metrics demonstrates a significant improvement in SSA scores for the Russian language, outperforming base models by 51%.