In the field of Natural Language Processing (NLP), Large-scale Language Models (LLMs) have demonstrated exceptional capabilities across a variety of tasks, including question answering, classification, and particularly natural language understanding. The integration of neural machine translation with LLMs presents significant potential, transforming the paradigms of cross-lingual communication and information exchange. This study investigates the foundational aspects of LLMs’ translation abilities and identifies effective training methodologies to equip them with multilingual capacities. We specifically explore the optimal timing for introducing translation capabilities to LLMs via supervised tasks, considering the inherent bilingual nature of machine translation. Key questions explored include whether it is more beneficial to integrate multiple languages during the pre-training or supervised fine-tuning (SFT) stages, how variations in language ratios influence LLMs’ translation abilities, and whether longer or shorter texts are more effective for training these models. This research conducts a thorough investigation by training multiple LLMs from scratch with parameter scales in the billions and enhances the robustness of our findings by upgrading the language capabilities of pre-trained open-source models with parameter scales reaching tens of billions. The aim is to provide a detailed analysis that elucidates the complexities of augmenting machine translation capabilities within LLMs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

\(E^3\) : Optimizing Language Model Training for Translation via Enhancing Efficiency and Effectiveness

  • Linqing Chen,
  • Weilei Wang,
  • Dongyang Hu

摘要

In the field of Natural Language Processing (NLP), Large-scale Language Models (LLMs) have demonstrated exceptional capabilities across a variety of tasks, including question answering, classification, and particularly natural language understanding. The integration of neural machine translation with LLMs presents significant potential, transforming the paradigms of cross-lingual communication and information exchange. This study investigates the foundational aspects of LLMs’ translation abilities and identifies effective training methodologies to equip them with multilingual capacities. We specifically explore the optimal timing for introducing translation capabilities to LLMs via supervised tasks, considering the inherent bilingual nature of machine translation. Key questions explored include whether it is more beneficial to integrate multiple languages during the pre-training or supervised fine-tuning (SFT) stages, how variations in language ratios influence LLMs’ translation abilities, and whether longer or shorter texts are more effective for training these models. This research conducts a thorough investigation by training multiple LLMs from scratch with parameter scales in the billions and enhances the robustness of our findings by upgrading the language capabilities of pre-trained open-source models with parameter scales reaching tens of billions. The aim is to provide a detailed analysis that elucidates the complexities of augmenting machine translation capabilities within LLMs.