Coarse-Grained Reconfigurable Arrays for High-Performance Low-Power Deep Neural Networks on Embedded Devices
摘要
Reconfigurable hardware architectures offer an effective platform for AI-powered Cyber-Physical Systems (CPS) by enabling efficient local computation, energy efficiency, and adaptability. This work presents a novel hardware architecture for executing Deep Neural Networks (DNNs) on resource-constrained edge devices. The architecture is organized as a Coarse-Grained Reconfigurable Array (CGRA) featuring flexible arithmetic units that support Precision Trade-off Floats (PT-Floats) II, customizable floating-point formats. By dynamically adjusting the precision of computations based on the sensitivity of different neural network layers, the proposed architecture achieves significant improvements in energy efficiency and memory footprint. Experimental results show 73% and 65% gains in die area and power consumption, respectively, to the cost of 0.02% in accuracy, compared with the standard IEEE-754. They evidence substantial performance gains in Convolutional Neural Network (CNN)-based image classification tasks while operating within the stringent power and resource constraints of edge devices.