Advanced Image Captioning Using Object Detectors and Large Language Models
摘要
Abstract
In this paper, a pipeline for composing titles to images using image-to-region-to-text and large language models is considered. By combining different datasets, the most adequate captioning model was trained and then its integration with the YandexGPT model was performed. The bilingual evaluation understudy and consensus-based image description evaluation metrics of our approach were improved several times compared to training on traditional datasets. In addition, a prompt for styling image captions in two styles (fun, normal) has been configured. The examples of advanced captioning generation using language models are given.