错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Vision-Language Models

  • Bharath Kumar Bolla,
  • Kalpa Subbaiah,
  • Sashi Kiran Kaata

摘要

In previous chapters, we optimized a model proficient in a single modality: text. We have fine-tuned, quantized, deployed, evaluated, and augmented it with external knowledge (RAG). However, real-world data is inherently multimodal. Human cognition processes visual, auditory, and textual information simultaneously to build a comprehensive understanding of the environment.