Enhancing Neural Machine Translation with Direct Preference Optimization Using Human Feedback
摘要
This paper presents a study on improving the quality of neural machine translation (NMT) for the English-Romanian language pair using Reinforcement Learning from Human Feedback (RLHF) via Direct Preference Optimization (DPO). Despite advancements in NMT, challenges remain, particularly for low-resource languages and personalized translations. By incorporating human feedback, the proposed approach demonstrates improvements in translation accuracy and naturalness. Although traditional metrics, such as BLEU and chrF++, yielded slightly lower scores for the DPO-trained model, human assessments indicate that the DPO-trained model better aligns with human preferences, particularly in everyday conversational contexts.