Exploring Vision-Language Models
摘要
In previous chapters, we optimized a model proficient in a single modality: text. We have fine-tuned, quantized, deployed, evaluated, and augmented it with external knowledge (RAG). However, real-world data is inherently multimodal. Human cognition processes visual, auditory, and textual information simultaneously to build a comprehensive understanding of the environment.