The rise of Vision-Language Models (VLMs) has advanced multimodal tasks but their high computational demands limit practical deployment on edge devices. To solve this, we propose TinyVision, a modular, distributed VLM architecture optimized for edge environments. TinyVision features a lightweight visual encoding mechanism that enhances feature extraction efficiency and decouples visual encoding from inference, reducing computational loads and improving privacy by minimizing raw data transmission. This architecture expands deployment possibilities to devices that were previously incompatible with traditional VLMs. Experimental results show that TinyVision achieves an average score of 44.8 on MM-series datasets, offering inference accuracy comparable to state-of-the-art VLMs while significantly reducing resource demands.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TinyVision: Distributed Vision-Language Model with Efficiency and Privacy for Edge Deployment

  • Shengfeng Lou,
  • Shangbao Ge,
  • Jianchao Yu,
  • Guoping Zhang

摘要

The rise of Vision-Language Models (VLMs) has advanced multimodal tasks but their high computational demands limit practical deployment on edge devices. To solve this, we propose TinyVision, a modular, distributed VLM architecture optimized for edge environments. TinyVision features a lightweight visual encoding mechanism that enhances feature extraction efficiency and decouples visual encoding from inference, reducing computational loads and improving privacy by minimizing raw data transmission. This architecture expands deployment possibilities to devices that were previously incompatible with traditional VLMs. Experimental results show that TinyVision achieves an average score of 44.8 on MM-series datasets, offering inference accuracy comparable to state-of-the-art VLMs while significantly reducing resource demands.