Exploring Efficient-Tuned Learning Audio Representation Method from BriVL
摘要
Recently, there has been an increase in the popularity of multimodal approaches in audio-related tasks, which involve using not only the audible modality but also textual or visual modalities in combination with sound. In this paper, we propose a robust audio representation learning method WavBriVL based on Bridging-Vision-and-Language (BriVL). It projects audio, image and text into a shared embedded space, so that multi-modal applications can be realized. We tested it on some downstream tasks and presented the images rearranged by our method and evaluated them qualitatively and quantitatively. The main purpose of this article is to: (1) Explore new correlation representations between audio and images; (2) Explore a new way to generate images using audio. The experimental results show that this method can effectively do a match on the audio image.