<p>Efficient inference of convolutional neural networks (CNNs) on energy-constrained devices has become increasingly important. Considering the extensive use of multiply-accumulate (MAC) operations in CNNs, efficient approximate multipliers have emerged as a promising solution. However, two major challenges remain: first, identifying a more suitable multiplier type for a given CNN model, and second, designing a reconfigurable multiplier architecture that supports a range of accuracy levels. This paper first investigates the key error metrics affecting classification accuracy in CNNs. Subsequently, three CNN-oriented 8-bit signed approximate multipliers are proposed, which rely on shift-based operations and a co-designed weight-precomputation scheme. Within this hardware-algorithm co-design, the precomputed weights are specifically optimized for the proposed multiplier architectures, thereby reducing effective CNN inference errors while simplifying the underlying hardware. These signed multipliers are appropriate for CNN inference applications. Furthermore, the proposed architecture is scalable and can be extended to support any <i>n</i>-bit multiplication. The proposed 8-bit structures are also applicable to operands of different bit widths, thereby offering flexibility that benefits CNN deployments. The proposed multipliers are evaluated across hardware and accuracy metrics during the inference phase of CNN models, including VGG10, VGG16, AlexNet, and ResNet-18. The proposed approximate multipliers were evaluated in both FPGA and ASIC implementations. The results demonstrate that the proposed multipliers achieve significant efficiency gains in both hardware implementation and accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Energy-efficient CNN acceleration using approximate multipliers with error-aware weight precomputation

  • Ladan Sayadi,
  • Mohammad Hossein Moaiyeri,
  • Somayeh Timarchi

摘要

Efficient inference of convolutional neural networks (CNNs) on energy-constrained devices has become increasingly important. Considering the extensive use of multiply-accumulate (MAC) operations in CNNs, efficient approximate multipliers have emerged as a promising solution. However, two major challenges remain: first, identifying a more suitable multiplier type for a given CNN model, and second, designing a reconfigurable multiplier architecture that supports a range of accuracy levels. This paper first investigates the key error metrics affecting classification accuracy in CNNs. Subsequently, three CNN-oriented 8-bit signed approximate multipliers are proposed, which rely on shift-based operations and a co-designed weight-precomputation scheme. Within this hardware-algorithm co-design, the precomputed weights are specifically optimized for the proposed multiplier architectures, thereby reducing effective CNN inference errors while simplifying the underlying hardware. These signed multipliers are appropriate for CNN inference applications. Furthermore, the proposed architecture is scalable and can be extended to support any n-bit multiplication. The proposed 8-bit structures are also applicable to operands of different bit widths, thereby offering flexibility that benefits CNN deployments. The proposed multipliers are evaluated across hardware and accuracy metrics during the inference phase of CNN models, including VGG10, VGG16, AlexNet, and ResNet-18. The proposed approximate multipliers were evaluated in both FPGA and ASIC implementations. The results demonstrate that the proposed multipliers achieve significant efficiency gains in both hardware implementation and accuracy.