Large language models (LLMs) have exhibited remarkable performance in various natural language processing tasks. In this paper, we describe our LLM-based machine translation system submitted to CCMT 2024 evaluation task. We investigated the effects of pre training data ratio, language quantity, instruction fine-tuning data ratio, and instruction construction method on the machine translation performance of large language models. We found that large language models have greater potential for multi-domain machine translation compared to traditional machine translation models. Within limits, increasing the number of newly added languages can enhance the overall translation capabilities for these new languages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Machine Translation Across Multiple Domains and Languages with Large Language Models

  • Hao Lu,
  • Rui Zhang,
  • Hui Huang,
  • Fuhai Song,
  • Junkai Liu,
  • Yican Ye,
  • Lang Lang,
  • Ziqing Zhao,
  • Muyun Yang,
  • Rui Cong

摘要

Large language models (LLMs) have exhibited remarkable performance in various natural language processing tasks. In this paper, we describe our LLM-based machine translation system submitted to CCMT 2024 evaluation task. We investigated the effects of pre training data ratio, language quantity, instruction fine-tuning data ratio, and instruction construction method on the machine translation performance of large language models. We found that large language models have greater potential for multi-domain machine translation compared to traditional machine translation models. Within limits, increasing the number of newly added languages can enhance the overall translation capabilities for these new languages.