<p>Achieving high-quality translation for literary works poses a unique challenge for machine translation models. This study compares Hungarian translations of Antoine de Saint-Exupéry’s novella, <i>The Little Prince</i>, produced by two leading neural machine translation (NMT) services (DeepL and Google Translate) and two large language models (LLMs) (Google Bard and ChatGPT 3.5). While the NMT tools achieved decent accuracy, their outputs often lacked the nuance required to capture the text’s literary essence. Notably, our research addresses a gap in prompt engineering by investigating whether the LLMs’ performance could be enhanced by using tailored, genre-specific prompts based on literary style guides, in contrast to baseline zero-shot outputs. Interestingly, this approach led to significant improvements for Google Bard in punctuation, grammar, and the preservation of literary devices. Conversely, the same prompt negatively affected the quality of the translation generated by ChatGPT 3.5. These findings suggest that while genre-specific prompts can guide certain LLMs toward self-correction, their effectiveness is highly model-dependent.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Taming AI for The Little Prince: a comparative analysis of NMTs and LLMs in Hungarian translation

  • Luyu Chen,
  • Lilla Varga,
  • Milad Mehdizadkhani

摘要

Achieving high-quality translation for literary works poses a unique challenge for machine translation models. This study compares Hungarian translations of Antoine de Saint-Exupéry’s novella, The Little Prince, produced by two leading neural machine translation (NMT) services (DeepL and Google Translate) and two large language models (LLMs) (Google Bard and ChatGPT 3.5). While the NMT tools achieved decent accuracy, their outputs often lacked the nuance required to capture the text’s literary essence. Notably, our research addresses a gap in prompt engineering by investigating whether the LLMs’ performance could be enhanced by using tailored, genre-specific prompts based on literary style guides, in contrast to baseline zero-shot outputs. Interestingly, this approach led to significant improvements for Google Bard in punctuation, grammar, and the preservation of literary devices. Conversely, the same prompt negatively affected the quality of the translation generated by ChatGPT 3.5. These findings suggest that while genre-specific prompts can guide certain LLMs toward self-correction, their effectiveness is highly model-dependent.