错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing high-vocabulary image annotation with a novel attention-based pooling

  • Ali Salar,
  • Ali Ahmadi

摘要

Given an image, we aim to assign a set of semantic labels to its visual content automatically. This is generally known as automatic image annotation (AIA). Images contain objects that can vary in size and position, with some only taking up a small region of the entire picture. The rise in the number of object classes also heightens this variety. Despite the achievement of promising results, the majority of the current methods have limited efficacy in the detection of small-scale objects. To make more effective use of spatial data compared to the global pooling method, we propose a modified transformer decoder layer that decreases computational complexity without sacrificing model performance. The study has conducted multiple experiments on four datasets, including three high-vocabulary small-scale datasets (Corel 5 k, IAPR TC-12, and Esp Game) and one large-scale dataset (Visual Genome) with a vocabulary list of 500 words. In comparison with existing state-of-the-art models, our approach achieves comparable results in F1-score, \({\text{N}}^{+}\) N + , and mean average precision (mAP) on small- and large-scale datasets.