<p>As large multimodal models (LMMs) advance rapidly across diverse multimodal understanding and generation tasks, the need for systematic and reliable evaluation frameworks becomes increasingly critical. To address this need, this survey provides a structured overview of LMM evaluation, centered around two main axes: multimodal evaluation for understanding and generation. (1) For understanding, a dual-perspective framework is introduced to distinguish benchmarks between general capabilities, which emphasize common tasks, and specialized capabilities, which reflect expert-level competence in domain-specific fields. (2) For generation, evaluation is organized by output modality, including image, video, audio, and 3D content. (3) From a community perspective, this survey further highlights authoritative leaderboards and foundational tools that have been instrumental in establishing a comprehensive evaluation ecosystem for LMMs. By unifying general-specialized understanding and modality-specific generation evaluations, this survey clarifies the current landscape and provides guidance for future research in the LMM evaluation field.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large multimodal models evaluation: a survey

  • Zicheng Zhang,
  • Junying Wang,
  • Farong Wen,
  • Yijin Guo,
  • Xiangyu Zhao,
  • Xinyu Fang,
  • Shengyuan Ding,
  • Ziheng Jia,
  • Jiahao Xiao,
  • Ye Shen,
  • Yushuo Zheng,
  • Xiaorong Zhu,
  • Yalun Wu,
  • Ziheng Jiao,
  • Wei Sun,
  • Zijian Chen,
  • Kaiwei Zhang,
  • Kang Fu,
  • Yuqin Cao,
  • Ming Hu,
  • Yue Zhou,
  • Xuemei Zhou,
  • Juntai Cao,
  • Wei Zhou,
  • Jinyu Cao,
  • Ronghui Li,
  • Donghao Zhou,
  • Yuan Tian,
  • Xiangyang Zhu,
  • Chunyi Li,
  • Haoning Wu,
  • Xiaohong Liu,
  • Junjun He,
  • Yu Zhou,
  • Hui Liu,
  • Lin Zhang,
  • Zesheng Wang,
  • Huiyu Duan,
  • Yingjie Zhou,
  • Xiongkuo Min,
  • Qi Jia,
  • Dongzhan Zhou,
  • Wenlong Zhang,
  • Jiezhang Cao,
  • Xue Yang,
  • Junzhi Yu,
  • Songyang Zhang,
  • Haodong Duan,
  • Guangtao Zhai

摘要

As large multimodal models (LMMs) advance rapidly across diverse multimodal understanding and generation tasks, the need for systematic and reliable evaluation frameworks becomes increasingly critical. To address this need, this survey provides a structured overview of LMM evaluation, centered around two main axes: multimodal evaluation for understanding and generation. (1) For understanding, a dual-perspective framework is introduced to distinguish benchmarks between general capabilities, which emphasize common tasks, and specialized capabilities, which reflect expert-level competence in domain-specific fields. (2) For generation, evaluation is organized by output modality, including image, video, audio, and 3D content. (3) From a community perspective, this survey further highlights authoritative leaderboards and foundational tools that have been instrumental in establishing a comprehensive evaluation ecosystem for LMMs. By unifying general-specialized understanding and modality-specific generation evaluations, this survey clarifies the current landscape and provides guidance for future research in the LMM evaluation field.