The emergence of ViT demonstrates the viability of Transformer in the field of computer vision, rivaling or even surpassing the performance of some CNN-based models. However, ViT’s substantial computation and memory demands present deployment challenges on resource-constrained devices. While various lightweight ViT techniques have been developed, they do not combine collaboration between heterogeneous devices for model splitting. Especially for large Transformer-based models, single-device deployment is infeasible, necessitating model splitting across multiple, even heterogeneous devices for distributed inference, and the optimal split point becomes crucial. Therefore, for a typical heterogeneous scenario of edge-cloud, we propose StressViT to split and compress visual Transformer through edge-cloud collaboration. StressViT considers the disparities in computility and memory between edge and cloud environments, striving to improve inference accuracy while satisfying memory and compression constraints. Evaluations on four heterogeneous devices and two datasets confirm the efficacy of StressViT, achieving nearly 50% reduction of end-to-end latency with almost no loss of accuracy. StressViT can also be applied to homogeneous devices for edge-edge or cloud-cloud splitting, and can be easily extended to split across more than two devices.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

StressViT: Splitting and Compressing Vision Transformer Through Edge-Cloud Collaboration

  • Changyao Lin,
  • Yi Liu,
  • Xiangyu Li,
  • Chengxiang Li,
  • Hao Zhang,
  • Jing Jin,
  • Jie Liu

摘要

The emergence of ViT demonstrates the viability of Transformer in the field of computer vision, rivaling or even surpassing the performance of some CNN-based models. However, ViT’s substantial computation and memory demands present deployment challenges on resource-constrained devices. While various lightweight ViT techniques have been developed, they do not combine collaboration between heterogeneous devices for model splitting. Especially for large Transformer-based models, single-device deployment is infeasible, necessitating model splitting across multiple, even heterogeneous devices for distributed inference, and the optimal split point becomes crucial. Therefore, for a typical heterogeneous scenario of edge-cloud, we propose StressViT to split and compress visual Transformer through edge-cloud collaboration. StressViT considers the disparities in computility and memory between edge and cloud environments, striving to improve inference accuracy while satisfying memory and compression constraints. Evaluations on four heterogeneous devices and two datasets confirm the efficacy of StressViT, achieving nearly 50% reduction of end-to-end latency with almost no loss of accuracy. StressViT can also be applied to homogeneous devices for edge-edge or cloud-cloud splitting, and can be easily extended to split across more than two devices.