错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Efficient-Tuned Learning Audio Representation Method from BriVL

  • Sen Fang,
  • Yangjian Wu,
  • Bowen Gao,
  • Jingwen Cai,
  • Teik Toe Teoh

摘要

Recently, there has been an increase in the popularity of multimodal approaches in audio-related tasks, which involve using not only the audible modality but also textual or visual modalities in combination with sound. In this paper, we propose a robust audio representation learning method WavBriVL based on Bridging-Vision-and-Language (BriVL). It projects audio, image and text into a shared embedded space, so that multi-modal applications can be realized. We tested it on some downstream tasks and presented the images rearranged by our method and evaluated them qualitatively and quantitatively. The main purpose of this article is to: (1) Explore new correlation representations between audio and images; (2) Explore a new way to generate images using audio. The experimental results show that this method can effectively do a match on the audio image.