Multimodal Machine Translation with Fusion of Generated Visual Information
摘要
Multi-modal machine translation requires a large amount of training data to improve performance. However, collecting and annotating labeled data typically requires a lot of time and human resources. For tasks such as natural language processing and computer vision, obtaining large-scale annotated data is a challenging problem that needs to be addressed. To address the issue of high cost in obtaining annotated data, this article proposes an effective method of integrating a pre-trained generative visual model for text-to-image generation into pure-text machine translation. To improve the modeling quality of semantic translation from text to image, we will perform conditional augmentation on the text used for image generation. The gate fusion mechanism is used to combine the generated image information with the conditionally augmented text information to assist in improving translation performance. To test the feasibility of the method, we conducted experiments on multiple different datasets. When generating text-to-image correspondence on the Multi30k English-to-German dataset, the Bleu score get a significant improvement in compared to pure-text translation. Experiments and visualization analysis show that the proposed method can effectively improve the quality of pure-text machine translation.