Multimodal Large Language Models (MLLMs) demonstrate strong reasoning capabilities by integrating visual and textual information with external knowledge, but their computational demands limit practical deployment. We propose a knowledge-guided structured pruning approach that leverages external knowledge graphs to inform compression decisions. Our method achieves a favorable trade-off between model size and performance: at 30% compression, we retain 89.2% of original accuracy while reducing inference time by 1.4x. Experiments on knowledge-grounded visual question answering show modest improvements over magnitude-based pruning baselines, though we observe increased hallucination rates typical of compressed models. Our approach provides a practical framework for deploying MLLMs in resource-constrained environments while maintaining reasonable performance on knowledge-intensive tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Knowledge-Guided Structured Pruning for Multimodal Language Models

  • P. Yadla,
  • L. Yadla

摘要

Multimodal Large Language Models (MLLMs) demonstrate strong reasoning capabilities by integrating visual and textual information with external knowledge, but their computational demands limit practical deployment. We propose a knowledge-guided structured pruning approach that leverages external knowledge graphs to inform compression decisions. Our method achieves a favorable trade-off between model size and performance: at 30% compression, we retain 89.2% of original accuracy while reducing inference time by 1.4x. Experiments on knowledge-grounded visual question answering show modest improvements over magnitude-based pruning baselines, though we observe increased hallucination rates typical of compressed models. Our approach provides a practical framework for deploying MLLMs in resource-constrained environments while maintaining reasonable performance on knowledge-intensive tasks.