Evaluating the limits of machine translation for poetry: a multidimensional framework
摘要
Automatic poetry translation remains a challenging task, as it requires not only semantic accuracy but also the preservation of stylistic and emotional elements. This study investigates the effectiveness of machine translation (MT) systems and large language models (LLMs) in poetry translation. Traditional automatic evaluation metrics often fail to capture the literary quality of such translations; therefore, we propose a three-phase evaluation framework that integrates complementary perspectives: (i) automatic metrics (BLEU, METEOR, and BERTScore) to assess lexical and semantic fidelity, (ii) topic modeling with BERTopic to perform an exploratory analysis of lexical-thematic alignment across translations, and (iii) expert human evaluation to examine poetic structure, style, fluency, and meaning preservation. This framework was applied to compare specialized MT systems (mBART, MarianMT, OpenNMT with RNN, and Google Translate) with LLMs such as ChatGPT and Maritaca AI across six language pairs involving English, French, and Portuguese (English–French, English–Portuguese, French–English, French–Portuguese, Portuguese–English, and Portuguese–French). Additionally, we investigate the impact of fine-tuning strategies using corpora of poems and song lyrics. The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment. However, human evaluation reveals that all systems struggle to replicate the poetic structure and stylistic nuances of the originals. The fine-tuning process did not produce improvements for all models; mBART showed notable gains, while the other models did not benefit significantly from domain adaptation.