<p>The automatic creation of textual descriptions from pictures, or image captioning, has advanced significantly in the last several years. It is a hybrid approach that employs natural language processing and computer vision. Several techniques are used in image-to-text synthesis, such as encoder–decoder frameworks, Unsupervised Learning, Reinforcement Learning, Transformer-based techniques, and attention processes. Despite these improvements, addressing image ambiguity and effectively collecting image context, emotions, and facts remains challenging. This work extends previous surveys by providing a comprehensive analysis encompassing the latest developments, evaluation challenges, and specific features like multilingual capabilities, narrative generation, detailed descriptions and potential integration with other applications. This work also investigates the evaluation metrics used in different application areas to train and assess image captioning algorithms. In addition, different issues and challenges, along with the primary outcomes from previous researchers, are covered to help researchers in future work, this gives researchers and developers the insights and tools they need to drive image captioning technology forward.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Holistic Review of Image-to-Text Conversion: Techniques, Evaluation Metrics, Multilingual Captioning, Storytelling and Integration

  • Anjali Sharma,
  • Mayank Aggarwal

摘要

The automatic creation of textual descriptions from pictures, or image captioning, has advanced significantly in the last several years. It is a hybrid approach that employs natural language processing and computer vision. Several techniques are used in image-to-text synthesis, such as encoder–decoder frameworks, Unsupervised Learning, Reinforcement Learning, Transformer-based techniques, and attention processes. Despite these improvements, addressing image ambiguity and effectively collecting image context, emotions, and facts remains challenging. This work extends previous surveys by providing a comprehensive analysis encompassing the latest developments, evaluation challenges, and specific features like multilingual capabilities, narrative generation, detailed descriptions and potential integration with other applications. This work also investigates the evaluation metrics used in different application areas to train and assess image captioning algorithms. In addition, different issues and challenges, along with the primary outcomes from previous researchers, are covered to help researchers in future work, this gives researchers and developers the insights and tools they need to drive image captioning technology forward.