This research article introduces novel approaches for realizing adaptive fixed-point Processing Engines (PEs) for DNN accelerators. We efficiently design a Multiply-Accumulate (MAC) unit and Activation Functions (AFs) that support computations on adaptive fixed-point represented as sfixed<N_f>, where ‘N’ (8 and 16 bits) and ‘f’ denote total width of data and fraction bits, respectively. We examine AF design: ROM-based versus pipelined Cordic for better hardware use and accuracy. The proposed adaptive fixed-point (N = 16) based MAC and AFs, implemented using ROM and Cordic, demonstrate negligible accuracy loss ( \(\le \) 1% for LeNet with MNIST, \(\le \) 1% for AlexNet with CIFAR-10 and \(\le \) 2% for VGG16 with CIFAR-10) when compared to the accuracy achieved with reference MAC and AFs realized based on Tensor with float32. Experimental results obtained from the Virtex-VCU118 Evaluation kit reveal the advantages of the ROM-based approach in cases of lower precision (sfixed&lt;8,f&gt;), showcasing a 76.8% reduction in LUT utilization and significant improvements in critical delay and maximum operating frequency. However, as precision increases to sfixed&lt;16,f&gt;, the Cordic-based approach exhibits a 76% reduction in LUT utilization and a 25.6% improvement in maximum frequency. These results illustrate the effective design of the proposed adaptive fixed-point PEs with MAC and AFs enabling DNN accelerators with flexible precision, which may be of requirement in various DNN models and layers.</N_f>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AFX-PE: Adaptive Fixed-Point Processing Engine for Neural Network Accelerators

  • Gopal Raut,
  • Ritambhara Thakur,
  • Pranose Edavoor,
  • David Selvakumar

摘要

This research article introduces novel approaches for realizing adaptive fixed-point Processing Engines (PEs) for DNN accelerators. We efficiently design a Multiply-Accumulate (MAC) unit and Activation Functions (AFs) that support computations on adaptive fixed-point represented as sfixed, where ‘N’ (8 and 16 bits) and ‘f’ denote total width of data and fraction bits, respectively. We examine AF design: ROM-based versus pipelined Cordic for better hardware use and accuracy. The proposed adaptive fixed-point (N = 16) based MAC and AFs, implemented using ROM and Cordic, demonstrate negligible accuracy loss ( \(\le \) 1% for LeNet with MNIST, \(\le \) 1% for AlexNet with CIFAR-10 and \(\le \) 2% for VGG16 with CIFAR-10) when compared to the accuracy achieved with reference MAC and AFs realized based on Tensor with float32. Experimental results obtained from the Virtex-VCU118 Evaluation kit reveal the advantages of the ROM-based approach in cases of lower precision (sfixed<8,f>), showcasing a 76.8% reduction in LUT utilization and significant improvements in critical delay and maximum operating frequency. However, as precision increases to sfixed<16,f>, the Cordic-based approach exhibits a 76% reduction in LUT utilization and a 25.6% improvement in maximum frequency. These results illustrate the effective design of the proposed adaptive fixed-point PEs with MAC and AFs enabling DNN accelerators with flexible precision, which may be of requirement in various DNN models and layers.