TinyVision: Distributed Vision-Language Model with Efficiency and Privacy for Edge Deployment
摘要
The rise of Vision-Language Models (VLMs) has advanced multimodal tasks but their high computational demands limit practical deployment on edge devices. To solve this, we propose TinyVision, a modular, distributed VLM architecture optimized for edge environments. TinyVision features a lightweight visual encoding mechanism that enhances feature extraction efficiency and decouples visual encoding from inference, reducing computational loads and improving privacy by minimizing raw data transmission. This architecture expands deployment possibilities to devices that were previously incompatible with traditional VLMs. Experimental results show that TinyVision achieves an average score of 44.8 on MM-series datasets, offering inference accuracy comparable to state-of-the-art VLMs while significantly reducing resource demands.