Most of the existing Surgical Visual Question Answering (VQA) systems use naive fusion strategies for text and image modalities and there is an absence of localized answering. The limited availability of annotated medical data and the complexity of domain-specific terminology have further limited the exploration of VQA systems for surgical procedures. We propose a Cross-Modal Attention (CroMA) based VQA system which can effectively fuse multimodal features from visual and textual sources. The fused embedding will feed a standard Class-Attention in Image Transformer (CaiT) module to the parallel classifier and the detector for joint prediction. Our experimental results on two public datasets suggest that CroMA based VQA system can better comprehend the surgical scene and localize the specific areas related to it with fewer parameters compared to other state-of-the-art (SOTA) models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CroMA: Cross-Modal Attention for Visual Question Answering in Robotic Surgery

  • Greetta Antonio,
  • Jobin Jose,
  • Sudhish N George,
  • Kiran Raja

摘要

Most of the existing Surgical Visual Question Answering (VQA) systems use naive fusion strategies for text and image modalities and there is an absence of localized answering. The limited availability of annotated medical data and the complexity of domain-specific terminology have further limited the exploration of VQA systems for surgical procedures. We propose a Cross-Modal Attention (CroMA) based VQA system which can effectively fuse multimodal features from visual and textual sources. The fused embedding will feed a standard Class-Attention in Image Transformer (CaiT) module to the parallel classifier and the detector for joint prediction. Our experimental results on two public datasets suggest that CroMA based VQA system can better comprehend the surgical scene and localize the specific areas related to it with fewer parameters compared to other state-of-the-art (SOTA) models.