Dual Transformer with Gated-Attention Fusion for News Disaster Image Captioning
摘要
The disaster news image captioning is designed to automatically analyze and comprehend the content of disaster news images, generating precise image description, aiming to alleviate the workload of news editors and expedite the dissemination of disaster information. Current image caption algorithms rely on objects and their relationships, while objects within the disaster news image class are extremely similar and easy to confuse. To address the above issues, we propose a Dual Transformer with Gated-attention Fusion (DTGF). On the encoder side, we propose Gated-attention to effectively interact with region features and grid features, and adaptively filter out semantic noise during the interaction. On the decoder side, we introduced two cross-attention designs, Concat and Parallel, aiming to integrate enhanced visual representation and further improve model performance. Experiments on the DNICC19k dataset show that our model has achieved state-of-the-art performance. It is capable of extracting refined visual information and generating more accurate and specific descriptions of disaster news images without relying on additional prior knowledge.