错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint Audio Captioning Transformer and Stable Diffusion for Audio-to-Image Generation

  • Jingtao Yu

摘要

In the present society where artificial intelligence (AI) is all over the place, there is a growing trend in using AI for innovative applications in various academic and industrial fields. This paper proposes a novel audio-to-image generation method that creates a pipeline connecting two models through deep learning. This approach has promising applications in the realms of audio analysis, natural language processing, and image generation and is expected to benefit digital art and audiovisual production. This paper introduces an audio subtitle converter model audio captioning transformer (ACT) for audio-to-text conversion and a stable diffusion-based model for text-to-image generation. Images are successfully generated from audio data by combining these models. The principle is to create a pipeline that feeds the text generated by ACT into stable diffusion to produce the final image. The study used the AudioCaps dataset, and the results of the experiment demonstrate the effectiveness of the methodology by showing clear, detailed images corresponding to the audio descriptions. This research opens up new areas for converting audio into images, making artificial intelligence an important tool for creative endeavors. Future work may involve domain-specific applications, such as generating artistic images from musical audio.