<p>Dialogue style transfer aims to generate responses that preserve semantic coherence while adhering to a specified stylistic persona. Recent studies have demonstrated that reinforcement learning combined with parameter-efficient fine-tuning can effectively optimize stylistic objectives; however, existing approaches are typically limited to a single character or rely on large, computationally intensive models. In this paper, we propose a unified reinforcement learning framework for multi-character dialogue style transfer that is scalable to small, distilled, and quantized language models. The proposed approach integrates Proximal Policy Optimization with Low-Rank Adaptation and a multi-class style reward model, enabling a single model to generate dialogue in multiple philosophical character styles via explicit conditioning. The method is evaluated across multiple model sizes, including GPT2-Large, GPT2, DistilGPT2, and INT8-quantized GPT2, and compared with untuned GPT2 baselines. Experimental results demonstrate that reinforcement learning substantially improves stylistic accuracy relative to untuned models, while smaller models retain most of the performance at significantly reduced computational cost. These findings indicate that the proposed framework is effective, scalable, and suitable for deployment in resource-constrained environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dialogue Style Transfer for Multiple Characters

  • R. Pravosud,
  • O. Marchenko

摘要

Dialogue style transfer aims to generate responses that preserve semantic coherence while adhering to a specified stylistic persona. Recent studies have demonstrated that reinforcement learning combined with parameter-efficient fine-tuning can effectively optimize stylistic objectives; however, existing approaches are typically limited to a single character or rely on large, computationally intensive models. In this paper, we propose a unified reinforcement learning framework for multi-character dialogue style transfer that is scalable to small, distilled, and quantized language models. The proposed approach integrates Proximal Policy Optimization with Low-Rank Adaptation and a multi-class style reward model, enabling a single model to generate dialogue in multiple philosophical character styles via explicit conditioning. The method is evaluated across multiple model sizes, including GPT2-Large, GPT2, DistilGPT2, and INT8-quantized GPT2, and compared with untuned GPT2 baselines. Experimental results demonstrate that reinforcement learning substantially improves stylistic accuracy relative to untuned models, while smaller models retain most of the performance at significantly reduced computational cost. These findings indicate that the proposed framework is effective, scalable, and suitable for deployment in resource-constrained environments.