<p>Multimodal graphs, which integrate diverse multimodal features and relations, are ubiquitous in real-world applications. However, existing multimodal graph learning methods are typically trained from scratch for specific graph data and tasks, failing to generalize across various multimodal graph data and tasks. To bridge this gap, we explore the potential of multimodal graph large language models (MG-LLM) to unify and generalize across diverse multimodal graph data and tasks. We propose a unified framework of multimodal graph data, tasks, and models, discovering the inherent multi-granularity and multi-scale characteristics in multimodal graphs. Specifically, we present five key desired characteristics for MG-LLM: (1) unified space for multimodal structures and attributes, (2) capability of handling diverse multimodal graph tasks, (3) multimodal graph in-context learning, (4) multimodal graph interaction with natural language, and (5) multimodal graph reasoning. We then elaborate on the key challenges, review existing literature, and highlight promising future research directions towards realizing these ambitious characteristics. Finally, we summarize existing multimodal graph datasets pertinent for model training. We believe this paper can contribute to the ongoing advancement of the research towards MG-LLM for generalization across multimodal graph data and tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards multimodal graph large language model

  • Xin Wang,
  • Zeyang Zhang,
  • Linxin Xiao,
  • Haibo Chen,
  • Chendi Ge,
  • Wenwu Zhu

摘要

Multimodal graphs, which integrate diverse multimodal features and relations, are ubiquitous in real-world applications. However, existing multimodal graph learning methods are typically trained from scratch for specific graph data and tasks, failing to generalize across various multimodal graph data and tasks. To bridge this gap, we explore the potential of multimodal graph large language models (MG-LLM) to unify and generalize across diverse multimodal graph data and tasks. We propose a unified framework of multimodal graph data, tasks, and models, discovering the inherent multi-granularity and multi-scale characteristics in multimodal graphs. Specifically, we present five key desired characteristics for MG-LLM: (1) unified space for multimodal structures and attributes, (2) capability of handling diverse multimodal graph tasks, (3) multimodal graph in-context learning, (4) multimodal graph interaction with natural language, and (5) multimodal graph reasoning. We then elaborate on the key challenges, review existing literature, and highlight promising future research directions towards realizing these ambitious characteristics. Finally, we summarize existing multimodal graph datasets pertinent for model training. We believe this paper can contribute to the ongoing advancement of the research towards MG-LLM for generalization across multimodal graph data and tasks.