<p>Deep neural networks (DNNs) have demonstrated remarkable performance across a wide range of applications. Despite their high accuracy, the large volume of parameters and high computational complexity pose significant challenges for deployment on resource-constrained platforms. Computing-in-memory (CIM) has emerged as a promising solution by integrating computing and memory units, thereby overcoming the traditional von Neumann bottleneck and improving overall efficiency. However, due to inherent limitations in device representation and data interface precision, CIM systems struggle to support high-precision computations. Consequently, model quantization becomes a key enabler for deploying DNNs on such platforms. This paper presents a comprehensive review of model quantization methods for CIM-based accelerators. First, we introduce the fundamental concepts of model quantization and CIM. Then, we review and analyze existing studies from three perspectives: fixed precision quantization, mixed precision quantization, and optimization of quantized models. Finally, we conclude with a discussion of current challenges and future directions in CIM-specific quantization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model quantization for computing-in-memory: a survey

  • Sifan Sun,
  • Jinyu Bai,
  • Hanting Chen,
  • Kaiwen Deng,
  • Zhiwei Xie,
  • Jingjing Li,
  • Bin Cao,
  • He Zhang,
  • Wang Kang,
  • Weisheng Zhao

摘要

Deep neural networks (DNNs) have demonstrated remarkable performance across a wide range of applications. Despite their high accuracy, the large volume of parameters and high computational complexity pose significant challenges for deployment on resource-constrained platforms. Computing-in-memory (CIM) has emerged as a promising solution by integrating computing and memory units, thereby overcoming the traditional von Neumann bottleneck and improving overall efficiency. However, due to inherent limitations in device representation and data interface precision, CIM systems struggle to support high-precision computations. Consequently, model quantization becomes a key enabler for deploying DNNs on such platforms. This paper presents a comprehensive review of model quantization methods for CIM-based accelerators. First, we introduce the fundamental concepts of model quantization and CIM. Then, we review and analyze existing studies from three perspectives: fixed precision quantization, mixed precision quantization, and optimization of quantized models. Finally, we conclude with a discussion of current challenges and future directions in CIM-specific quantization.