Cross-lingual image captioning is pivotal in global communication, enabling seamless comprehension and dissemination of visual content across diverse languages. Existing approaches predominantly focus on learning symmetric feature spaces between modalities and languages, often neglecting the discriminative relationships inherent in cross-modal, cross-lingual sample pairs. However, cross-lingual, cross-modal tasks inherently require asymmetric feature learning tailored to specific target languages, while also demanding substantial datasets-exposing the limitations of current techniques. This paper introduces Asymmetric Cross-modal Cross-lingual Feature Learning (ACCFL), a novel method tailored for cross-modality feature alignment under constrained data availability. ACCFL addresses the challenges of non-symmetric feature learning and data scarcity by leveraging vision-language pre-trained models and effectively incorporating metric learning to mine hard negative samples. This approach enhances the precision and language-specificity of cross-modal feature alignment in image captioning scenarios. Experimental results demonstrate the superiority of the proposed method, yielding more accurate and robust feature representations for natural image captions in target languages, surpassing the performance of traditional symmetric methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Asymmetric Language-Aware Feature Learning for Low-Resource Cross-Lingual Image Caption

  • Miaomiao Lyu,
  • Junjie Chen,
  • Meishan Zhang

摘要

Cross-lingual image captioning is pivotal in global communication, enabling seamless comprehension and dissemination of visual content across diverse languages. Existing approaches predominantly focus on learning symmetric feature spaces between modalities and languages, often neglecting the discriminative relationships inherent in cross-modal, cross-lingual sample pairs. However, cross-lingual, cross-modal tasks inherently require asymmetric feature learning tailored to specific target languages, while also demanding substantial datasets-exposing the limitations of current techniques. This paper introduces Asymmetric Cross-modal Cross-lingual Feature Learning (ACCFL), a novel method tailored for cross-modality feature alignment under constrained data availability. ACCFL addresses the challenges of non-symmetric feature learning and data scarcity by leveraging vision-language pre-trained models and effectively incorporating metric learning to mine hard negative samples. This approach enhances the precision and language-specificity of cross-modal feature alignment in image captioning scenarios. Experimental results demonstrate the superiority of the proposed method, yielding more accurate and robust feature representations for natural image captions in target languages, surpassing the performance of traditional symmetric methods.