The connection between vision and language plays a crucial role in generative intelligence. Consequently, a significant amount of research has focused on image captioning, which involves describing images using grammatically and semantically meaningful sentences. Since 2015, the image captioning task has typically been addressed through a pipeline comprising a visual encoder and a language model for text generation. In recent years, with the development of large language models (LLMs), the two core components of image captioning—visual encoding and text generation—have seen significant improvements through the use of object regions, attributes, multimodal connections, full attention methods, and BERT-like early fusion strategies. Despite remarkable progress, the field of LLM-based image captioning remains inconclusive. This study overviews LLM-based image captioning methods, covering various aspects from visual encoding and text generation to training strategies, datasets, and multimodality. We quantitatively compare many of the latest relevant technical approaches to identify the most impactful innovations in architecture and training strategies. Additionally, this paper discusses the numerous variants of this problem and the unresolved challenges they present. Ultimately, this study aims to serve as a tool for understanding the existing literature and highlighting research directions that promise the best synergy between computer vision and natural language processing in the future.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Overview of Image Captioning Generation Based on Large Language Models

  • Jiahao Zhang,
  • Dehai Zhang,
  • Xu He,
  • Feng Gao,
  • Xiaohan Li

摘要

The connection between vision and language plays a crucial role in generative intelligence. Consequently, a significant amount of research has focused on image captioning, which involves describing images using grammatically and semantically meaningful sentences. Since 2015, the image captioning task has typically been addressed through a pipeline comprising a visual encoder and a language model for text generation. In recent years, with the development of large language models (LLMs), the two core components of image captioning—visual encoding and text generation—have seen significant improvements through the use of object regions, attributes, multimodal connections, full attention methods, and BERT-like early fusion strategies. Despite remarkable progress, the field of LLM-based image captioning remains inconclusive. This study overviews LLM-based image captioning methods, covering various aspects from visual encoding and text generation to training strategies, datasets, and multimodality. We quantitatively compare many of the latest relevant technical approaches to identify the most impactful innovations in architecture and training strategies. Additionally, this paper discusses the numerous variants of this problem and the unresolved challenges they present. Ultimately, this study aims to serve as a tool for understanding the existing literature and highlighting research directions that promise the best synergy between computer vision and natural language processing in the future.