Image captioning is a crucial task in the intersection of computer vision (CV) and natural language processing (NLP), enabling the generation of descriptive sentences for images. While significant progress has been made in English-based image captioning, challenges remain for languages with complex morphology and syntax, such as Arabic, due to the need for extensive linguistic resources and handling rich inflectional forms. In this work, we introduce a novel image captioning model that leverages the Swin-Transformer architecture to enhance the performance of a Pure Transformer (PureT) model. The Swin-Transformer, with its hierarchical structure and shifted windows mechanism, effectively captures both local and global image features, thereby improving the captioning quality. Our model was evaluated using standard metrics including BLEU, METEOR, ROUGE-L, and CIDEr, demonstrating significant improvements over existing models and achieving higher accuracy and better quality in generating Arabic captions. Specifically, our model achieved BLEU-1 of 76.4, BLEU-4 of 32.2, METEOR of 37.9, ROUGE-L of 54.2, and CIDEr of 109.6 for the uncleaned Arabic corpus, showcasing its superior performance in both accuracy and quality of captions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Arabic Image Captioning Using Transformer-Based Encoder-Decoder Model

  • Yasser Alhabashi,
  • Lahouari Ghouti,
  • Serry Sibaee,
  • Anis Koubaa

摘要

Image captioning is a crucial task in the intersection of computer vision (CV) and natural language processing (NLP), enabling the generation of descriptive sentences for images. While significant progress has been made in English-based image captioning, challenges remain for languages with complex morphology and syntax, such as Arabic, due to the need for extensive linguistic resources and handling rich inflectional forms. In this work, we introduce a novel image captioning model that leverages the Swin-Transformer architecture to enhance the performance of a Pure Transformer (PureT) model. The Swin-Transformer, with its hierarchical structure and shifted windows mechanism, effectively captures both local and global image features, thereby improving the captioning quality. Our model was evaluated using standard metrics including BLEU, METEOR, ROUGE-L, and CIDEr, demonstrating significant improvements over existing models and achieving higher accuracy and better quality in generating Arabic captions. Specifically, our model achieved BLEU-1 of 76.4, BLEU-4 of 32.2, METEOR of 37.9, ROUGE-L of 54.2, and CIDEr of 109.6 for the uncleaned Arabic corpus, showcasing its superior performance in both accuracy and quality of captions.