错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Chinese‒Korean Cross-Language Transfer Method That is Based on Language Features

  • Jia Meng,
  • Guozhe Jin,
  • Yahui Zhao,
  • Rongyi Cui

摘要

In view of the current problem of insufficient utilization of language characteristics in the field of cross-language transfer between China and South Korea, a method of cross-language transfer between China and South Korea based on linguistic features of Chinese characters is proposed. This method first uses Hanjaro to process the Chinese characters in the Korean sentence into the traditional Chinese form. Second, the original, unprocessed Korean sentence is constructed into a positive example sentence via random dropout. These two sentences are used as a positive example pair for contrastive learning training. Moreover, the preprocessed Korean sentences and the processed Korean sentences are used as a set of parallel sentences for confrontation training. The training goal is to enable the pretrained model encoder to generate parallel sentences. It can confuse the sentence embedding of the discriminator, thereby bringing the sentence vector closer. Finally, we use our model and the original multilingual pre-trained language model to train on the Chinese dataset and conduct experiments on the Korean dataset in a zero-shot manner to obtain the transfer results. This paper uses the public dataset PAWS-X to conduct the experiments. The experimental findings demonstrate that our model enhances the cross-language transfer capability from Chinese to Korean in comparison with the original multilingual pre-trained model.