Efficient Processing Element Architecture Using Hybrid Approximate Multipliers and Parallel Prefix Adders for CNN Accelerators
摘要
Convolutional Neural Networks (CNNs) have become highly accurate and are extensively used for image identification. However, implementing CNNs in hardware creates a substantial challenge due to the rise of deep learning applications. Therefore, optimizing hardware design for efficient CNN acceleration is crucial. A vital element of CNN accelerator architecture is the processing element (PE), which is responsible for the convolution operation. This paper implements the PE block with different adders and multipliers to reduce hardware utilization and power consumption. In this paper, we replaced bulky MAC units and conventional adder trees with a hybrid (Modified Booth Encoding (MBE) multiplier and WALLACE tree (WT)) and approximate parallel prefix adders (PPA) to reduce the area and delay. All the designs are implemented and synthesized on Xilinx Zynq FPGA. The design was compared with conventional approaches regarding power consumption, area, and delay. The outcome showed that the proposed design outperformed the conventional design, reducing area, delay, and power consumption.