Large Language Model on Multi-Modal Data
摘要
Large Language Models (LLMs) are used as the brain to do multimodal tasks in the emerging field of multimodal Large Language Model (MLLM) represented by GPT-4V. Surprisingly, MLLM’s emergent capabilities—like OCR-free math reasoning and the ability to write stories based on images—are uncommon in conventional multimodal approaches and point to a possible route towards artificial general intelligence. In an attempt to create MLLMs that are on scale with or even superior to GPT-4V, researchers in academia and industry have been working surprisingly quickly to push the boundaries of their field. This paper explores the potential of multimodal large language models, the core concepts of LLM, and the distinct design that sets the existing multimodal LLM apart. Additionally, it has outlined the difficulties and restrictions that the existing LLM faces, including model complexity and data heterogeneity.