<p>In-memory computing (IMC) has emerged as a promising approach for accelerating deep neural network (DNN) inference by relocating computations to memory arrays. However, the efficacy of analog IMC diminishes when higher computational precision is required due to inherent device non-idealities. In this paper, we present a reconfigurable heterogeneous architecture that integrates a digital computing unit (DCU) with an analog IMC unit (AIMCU). The computational data is partitioned into most significant bits (MSBs) and least significant bits (LSBs); the sparse MSBs are processed by the DCU with lossless precision, and the dense LSBs are computed by the AIMCU for high energy efficiency, thereby enhancing inference accuracy and optimizing area efficiency. The architecture also features multiple modes that support variable-precision input splitting and weight splitting computation. Additionally, by leveraging hardware characteristics, we have developed several optimization strategies for neural network deployment, including parameter splitting, shifting algorithms, and sparse weight mapping. The experimental results show that the perceptual evaluation of speech quality (PESQ) of the deep complex convolution recurrent network (DCCRN) improved by 28.98%, while the peak signal-to-noise ratio (PSNR) of the super-resolution network (SRN) increased by 17.27%. Compared to previous state-of-the-art (SOTA) work, the reconfigurable heterogeneous-IMC-based&#xa0;system on a chip (SoC) demonstrates a significant improvement in energy efficiency while achieving accuracy close to that of pure digital computing.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A reconfigurable heterogeneous in-memory computing architecture for variable precision computation: a software-hardware co-design approach

  • Yizhe Chen,
  • Hanjie Liu,
  • Saiya Wang,
  • Jinyao Mi,
  • Xiaodi Xing,
  • Yuexi Lv,
  • Aifei Zhang,
  • Lichuan Luo,
  • Yong Pei,
  • Minghua Tang,
  • Wang Kang

摘要

In-memory computing (IMC) has emerged as a promising approach for accelerating deep neural network (DNN) inference by relocating computations to memory arrays. However, the efficacy of analog IMC diminishes when higher computational precision is required due to inherent device non-idealities. In this paper, we present a reconfigurable heterogeneous architecture that integrates a digital computing unit (DCU) with an analog IMC unit (AIMCU). The computational data is partitioned into most significant bits (MSBs) and least significant bits (LSBs); the sparse MSBs are processed by the DCU with lossless precision, and the dense LSBs are computed by the AIMCU for high energy efficiency, thereby enhancing inference accuracy and optimizing area efficiency. The architecture also features multiple modes that support variable-precision input splitting and weight splitting computation. Additionally, by leveraging hardware characteristics, we have developed several optimization strategies for neural network deployment, including parameter splitting, shifting algorithms, and sparse weight mapping. The experimental results show that the perceptual evaluation of speech quality (PESQ) of the deep complex convolution recurrent network (DCCRN) improved by 28.98%, while the peak signal-to-noise ratio (PSNR) of the super-resolution network (SRN) increased by 17.27%. Compared to previous state-of-the-art (SOTA) work, the reconfigurable heterogeneous-IMC-based system on a chip (SoC) demonstrates a significant improvement in energy efficiency while achieving accuracy close to that of pure digital computing.