This chapter explores the evolution and applications of large language models (LLMs) in natural language processing, detailing their architecture, training methodologies, and usage in tasks like information retrieval and text generation. It then examines 2D visual language models (2D VLMs), which integrate visual and textual data for applications such as image captioning and visual question-answering, with insights into models like CLIP and BLIP. The chapter progresses to 2D multi-modal large language models (2D MLLMs), highlighting their enhanced contextual understanding, with examples like Flamingo, BLIP-2, and LLaVA. It further delves into 3D MLLMs, which process 3D data to understand and interact with 3D scenes and objects. Additionally, the concept of embodied AI is introduced, demonstrating the integration of perception, cognition, and action for complex tasks, exemplified by Google’s PaLM-E and DeepMind’s RT-2. The chapter concludes by anticipating future advancements in AI, particularly in robotics and advanced task automation, driven by the ongoing development of 3D MLLMs and embodied AI.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Point Cloud-Language Multi-modal Learning

  • Wei Gao,
  • Ge Li

摘要

This chapter explores the evolution and applications of large language models (LLMs) in natural language processing, detailing their architecture, training methodologies, and usage in tasks like information retrieval and text generation. It then examines 2D visual language models (2D VLMs), which integrate visual and textual data for applications such as image captioning and visual question-answering, with insights into models like CLIP and BLIP. The chapter progresses to 2D multi-modal large language models (2D MLLMs), highlighting their enhanced contextual understanding, with examples like Flamingo, BLIP-2, and LLaVA. It further delves into 3D MLLMs, which process 3D data to understand and interact with 3D scenes and objects. Additionally, the concept of embodied AI is introduced, demonstrating the integration of perception, cognition, and action for complex tasks, exemplified by Google’s PaLM-E and DeepMind’s RT-2. The chapter concludes by anticipating future advancements in AI, particularly in robotics and advanced task automation, driven by the ongoing development of 3D MLLMs and embodied AI.