<p>Traditional computer vision generally solves each single task independently by a specialist model with the task instruction implicitly considered and designed in the model architecture. This simply leads to two constraints in: (1) task-specific models where each model is trained for one specific task, hindering its scalability and synergy across diverse tasks; (2) pre-defined and fixed model interfaces that have limited interactivity and adaptability in following user’s task instructions. Visual Instruction Tuning (VIT), which learns from a wide range of vision tasks as described by natural language instructions, has recently been intensively studied to mitigate the constraints of specialist models. It fine-tunes a large vision model with natural language as general task instructions, aiming for a general-purpose multimodal large language model (MLLM) that can follow various language instructions and potentially solve various user-specified vision tasks. This work aims to provide a systematic and comprehensive review of visual instruction tuning that covers six key aspects including: (1) the background of vision task paradigm and its development towards VIT; (2) the foundations of VIT including commonly used network architectures, visual instruction tuning frameworks and objectives, as well as evaluation setups and tasks; (3) widely adopted benchmarks in visual instruction tuning and evaluations; (4) a thorough review of existing VIT techniques as categorized by both vision tasks and method designs, highlighting their major contributions, strengths, as well as constraints; (5) comparison and discussion of VIT methods over various instruction-following benchmarks; (6) challenges, possible research directions and research topics in the future visual instruction tuning study. A project associated with this work has been created at &#xa0;<a href="https://github.com/jingyi0000/Awesome-Visual-Instruction-Tuning">[link]</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual Instruction Tuning towards General-Purpose Multimodal Large Language Model: A Survey

  • Jiaxing Huang,
  • Jingyi Zhang,
  • Kai Jiang,
  • Han Qiu,
  • Xiaoqin Zhang,
  • Ling Shao,
  • Shijian Lu,
  • Dacheng Tao

摘要

Traditional computer vision generally solves each single task independently by a specialist model with the task instruction implicitly considered and designed in the model architecture. This simply leads to two constraints in: (1) task-specific models where each model is trained for one specific task, hindering its scalability and synergy across diverse tasks; (2) pre-defined and fixed model interfaces that have limited interactivity and adaptability in following user’s task instructions. Visual Instruction Tuning (VIT), which learns from a wide range of vision tasks as described by natural language instructions, has recently been intensively studied to mitigate the constraints of specialist models. It fine-tunes a large vision model with natural language as general task instructions, aiming for a general-purpose multimodal large language model (MLLM) that can follow various language instructions and potentially solve various user-specified vision tasks. This work aims to provide a systematic and comprehensive review of visual instruction tuning that covers six key aspects including: (1) the background of vision task paradigm and its development towards VIT; (2) the foundations of VIT including commonly used network architectures, visual instruction tuning frameworks and objectives, as well as evaluation setups and tasks; (3) widely adopted benchmarks in visual instruction tuning and evaluations; (4) a thorough review of existing VIT techniques as categorized by both vision tasks and method designs, highlighting their major contributions, strengths, as well as constraints; (5) comparison and discussion of VIT methods over various instruction-following benchmarks; (6) challenges, possible research directions and research topics in the future visual instruction tuning study. A project associated with this work has been created at  [link].