Enhancing Machine Translation Across Multiple Domains and Languages with Large Language Models
摘要
Large language models (LLMs) have exhibited remarkable performance in various natural language processing tasks. In this paper, we describe our LLM-based machine translation system submitted to CCMT 2024 evaluation task. We investigated the effects of pre training data ratio, language quantity, instruction fine-tuning data ratio, and instruction construction method on the machine translation performance of large language models. We found that large language models have greater potential for multi-domain machine translation compared to traditional machine translation models. Within limits, increasing the number of newly added languages can enhance the overall translation capabilities for these new languages.