<p>The increasing use of social media has posed challenges for search and raised concerns about effective classification approaches for this massive amount of data. To manage this influx of information effectively, leveraging hashtags emerges as a valuable option for enhancing information diffusion. This paper introduces TAMTE, a topic-attentive multimodal transformer encoder for hashtag recommendation task. Specifically, introduce a topic leveraging self-attention mechanism for cross-modal representation learning. The proposed model first encodes the text and image inputs separately using dedicated encoders, and a joint embedding is produced from image and text embeddings. Then, the microblog caption text and the image text generated from the image-to-text model are combined and fed into a topic modeling module to capture latent themes. Finally, the proposed approach proceeds using topic-attentive self-attentive mechanism between the topic embedding and the joint feature embedding. This facilitates more contextually relevant hashtag recommendations. Experimental results show the superiority of the proposed TAMTE method in terms of precision, recall, and F1-score compared to the existing state-of-the-art baseline methods. By effectively utilizing self-attention across multimodal and topic word representations, the proposed approach offers an effective solution for hashtag recommendation in social media contexts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TAMTE: topic-attentive multimodal transformer encoder for hashtag recommendation

  • Mian Muhammad Yasir Khalil,
  • Bo Chen,
  • Muhammad Arslan Rauf,
  • Weidong Wang,
  • Qingxian Wang

摘要

The increasing use of social media has posed challenges for search and raised concerns about effective classification approaches for this massive amount of data. To manage this influx of information effectively, leveraging hashtags emerges as a valuable option for enhancing information diffusion. This paper introduces TAMTE, a topic-attentive multimodal transformer encoder for hashtag recommendation task. Specifically, introduce a topic leveraging self-attention mechanism for cross-modal representation learning. The proposed model first encodes the text and image inputs separately using dedicated encoders, and a joint embedding is produced from image and text embeddings. Then, the microblog caption text and the image text generated from the image-to-text model are combined and fed into a topic modeling module to capture latent themes. Finally, the proposed approach proceeds using topic-attentive self-attentive mechanism between the topic embedding and the joint feature embedding. This facilitates more contextually relevant hashtag recommendations. Experimental results show the superiority of the proposed TAMTE method in terms of precision, recall, and F1-score compared to the existing state-of-the-art baseline methods. By effectively utilizing self-attention across multimodal and topic word representations, the proposed approach offers an effective solution for hashtag recommendation in social media contexts.