Multimodal context-aware translation of the endangered dongba script
摘要
Dongba script is one of the few pictographic writing systems still in use, characterized by nonlinear structure and strong integration of text and images, which hinders direct application of sequence-based machine translation. We make the first attempt to introduce an open-source multimodal large language model (MLLM) for paragraph-level Dongba-to-Chinese translation. To handle its nonlinear semantics, we propose a structured contextual semantic augmentation (SCSA) method, which segments Dongba images into coherent units and generates diverse training samples through reordering and spatial splicing. This guides the model to capture independent semantic units and their combinatorial relations. To support this task, we construct DongbaMMTCorpus, the first large-scale Dongba-Chinese image-text parallel corpus. Experiments indicate that SCSA can improve the ability of MLLMs to process nonlinear structures and enhance translation quality, providing a preliminary step toward automated translation of pictographic scripts.