Multimodal Enhanced Target Representation for Machine Translation
摘要
Recently, visual annotation is successfully introduced into the encoder-decoder framework to open a new multimodal neural machine translation (MMT) scenario. However, it is still difficult for MMT to simulate the ground-truth target future context due to the auto-regressive decoder. The visual annotation, which presents the content described in the bilingual parallel sentence pair, provides a ground-truth global contextual information for the auto-regressive decoder. Thus, we propose to directly learn a practical target for future contextual information from the visual modality space, thereby enhancing the target representation of the decoder. Empirical results on several widely used multimodal translation benchmarks demonstrated that the proposed approach significantly improved the performance of MMT over strong baselines.