Joint Audio Captioning Transformer and Stable Diffusion for Audio-to-Image Generation
摘要
In the present society where artificial intelligence (AI) is all over the place, there is a growing trend in using AI for innovative applications in various academic and industrial fields. This paper proposes a novel audio-to-image generation method that creates a pipeline connecting two models through deep learning. This approach has promising applications in the realms of audio analysis, natural language processing, and image generation and is expected to benefit digital art and audiovisual production. This paper introduces an audio subtitle converter model audio captioning transformer (ACT) for audio-to-text conversion and a stable diffusion-based model for text-to-image generation. Images are successfully generated from audio data by combining these models. The principle is to create a pipeline that feeds the text generated by ACT into stable diffusion to produce the final image. The study used the AudioCaps dataset, and the results of the experiment demonstrate the effectiveness of the methodology by showing clear, detailed images corresponding to the audio descriptions. This research opens up new areas for converting audio into images, making artificial intelligence an important tool for creative endeavors. Future work may involve domain-specific applications, such as generating artistic images from musical audio.