Image Caption Extending Using LLM and Style Transfer
摘要
This paper discusses a pipeline for generating captions for images using image-to-text models and large language models. By combining various datasets, the most suitable captioning model was trained, and this model was subsequently integrated with YandexGPT. Our approach demonstrated significant improvements in BLEU and CIDEr scores compared to traditional dataset training. Additionally, prompts for styling captions in three different tones (fun, angry and normal) were configured. Several intriguing examples of caption generation using language models are provided. In addition, a large language model was trained to perform prompt styling for subsequent querying of the generative image model. This approach showed the possibility of transferring the gist part of the image to another style.