Visual knowledge is primarily acquired through visual perception, but it is often exclusively represented in natural language, neglecting the collaborative nature of multisensory perception. To address this limitation, this paper proposes an audio-guided approach to visual knowledge representation. By integrating auditory cues into visual captioning, the model enhances environmental understanding through multisensory collaboration. Furthermore, the introduction of an audio-visual multimodal mutual information graph enriches the semantic content of visual captions. Additionally, while research on multimodal perception data is extensive, audio-visual datasets often lack fine-grained annotations. To address this issue, we construct a fine-grained multimodal dataset. Finally, experimental validation through multimodal-guided visual captioning and link prediction tasks demonstrates the effectiveness of this approach compared to existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio-Guided Visual Knowledge Representation

  • Fei Yu,
  • Zhiguo Wan,
  • Yuehua Li

摘要

Visual knowledge is primarily acquired through visual perception, but it is often exclusively represented in natural language, neglecting the collaborative nature of multisensory perception. To address this limitation, this paper proposes an audio-guided approach to visual knowledge representation. By integrating auditory cues into visual captioning, the model enhances environmental understanding through multisensory collaboration. Furthermore, the introduction of an audio-visual multimodal mutual information graph enriches the semantic content of visual captions. Additionally, while research on multimodal perception data is extensive, audio-visual datasets often lack fine-grained annotations. To address this issue, we construct a fine-grained multimodal dataset. Finally, experimental validation through multimodal-guided visual captioning and link prediction tasks demonstrates the effectiveness of this approach compared to existing methods.