<p>Dongba script is one of the few pictographic writing systems still in use, characterized by nonlinear structure and strong integration of text and images, which hinders direct application of sequence-based machine translation. We make the first attempt to introduce an open-source multimodal large language model (MLLM) for paragraph-level Dongba-to-Chinese translation. To handle its nonlinear semantics, we propose a structured contextual semantic augmentation (SCSA) method, which segments Dongba images into coherent units and generates diverse training samples through reordering and spatial splicing. This guides the model to capture independent semantic units and their combinatorial relations. To support this task, we construct DongbaMMTCorpus, the first large-scale Dongba-Chinese image-text parallel corpus. Experiments indicate that SCSA can improve the ability of MLLMs to process nonlinear structures and enhance translation quality, providing a preliminary step toward automated translation of pictographic scripts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal context-aware translation of the endangered dongba script

  • Shuo Li,
  • Xiaojun Bi,
  • Junyao Xing,
  • Ziyue Wang,
  • Fuwen Luo,
  • Peng Li,
  • Yang Liu

摘要

Dongba script is one of the few pictographic writing systems still in use, characterized by nonlinear structure and strong integration of text and images, which hinders direct application of sequence-based machine translation. We make the first attempt to introduce an open-source multimodal large language model (MLLM) for paragraph-level Dongba-to-Chinese translation. To handle its nonlinear semantics, we propose a structured contextual semantic augmentation (SCSA) method, which segments Dongba images into coherent units and generates diverse training samples through reordering and spatial splicing. This guides the model to capture independent semantic units and their combinatorial relations. To support this task, we construct DongbaMMTCorpus, the first large-scale Dongba-Chinese image-text parallel corpus. Experiments indicate that SCSA can improve the ability of MLLMs to process nonlinear structures and enhance translation quality, providing a preliminary step toward automated translation of pictographic scripts.