When exploring the development of Artificial General Intelligence (AGI), a critical task for these models involves interpreting and processing information from multiple image inputs. However, Large Multimodal Models (LMMs) encounter two issues in such scenarios: (1) a lack of fine-grained perception, and (2) a tendency to blend information across multiple images. To better investigate the capability of LMMs to perceive fine-grained visual details when dealing with multiple input images, we built a benchmark for evaluating LMM with multiple image inputs - MIMU(Muti-Image Inputs Multimodal Understanding Benchmark). The benchmark focuses on two scenarios: first, image-to-image matching (to evaluate whether LMMs can effectively reason and pair relevant images), and second, multi-image-to-text matching (to assess whether LMMs can accurately capture and summarize detailed image information). We conduct evaluations on a range of both open-source and closed-source large models, including GPT-4V, Gemini, OpenFlamingo, and MMICL. Although GPT-4V achieves the best results in all metrics, it still has a significant gap from Human Evaluation. To enhance model performance, we further develop a Contrastive Chain-of-Thought (CoCoT) prompting approach based on multi-input multimodal models. This method requires LMMs to compare the similarities and differences among multiple image inputs, and then guide the models to answer detailed questions about multi-image inputs based on the identified similarities and differences. Our experimental results showcase CoCoT’s proficiency in enhancing the multi-image comprehension capabilities of large multimodal models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Benchmark and Chain-of-Thought Prompting Strategy for Large Multimodal Models with Multiple Image Inputs

  • Daoan Zhang,
  • Junming Yang,
  • Hanjia Lyu,
  • Zijian Jin,
  • Yuan Yao,
  • Mingkai Chen,
  • Jiebo Luo

摘要

When exploring the development of Artificial General Intelligence (AGI), a critical task for these models involves interpreting and processing information from multiple image inputs. However, Large Multimodal Models (LMMs) encounter two issues in such scenarios: (1) a lack of fine-grained perception, and (2) a tendency to blend information across multiple images. To better investigate the capability of LMMs to perceive fine-grained visual details when dealing with multiple input images, we built a benchmark for evaluating LMM with multiple image inputs - MIMU(Muti-Image Inputs Multimodal Understanding Benchmark). The benchmark focuses on two scenarios: first, image-to-image matching (to evaluate whether LMMs can effectively reason and pair relevant images), and second, multi-image-to-text matching (to assess whether LMMs can accurately capture and summarize detailed image information). We conduct evaluations on a range of both open-source and closed-source large models, including GPT-4V, Gemini, OpenFlamingo, and MMICL. Although GPT-4V achieves the best results in all metrics, it still has a significant gap from Human Evaluation. To enhance model performance, we further develop a Contrastive Chain-of-Thought (CoCoT) prompting approach based on multi-input multimodal models. This method requires LMMs to compare the similarities and differences among multiple image inputs, and then guide the models to answer detailed questions about multi-image inputs based on the identified similarities and differences. Our experimental results showcase CoCoT’s proficiency in enhancing the multi-image comprehension capabilities of large multimodal models.