<p>This work proposes a dual-sparsity-aware processing element (PE) incorporated within the systolic array network, which can avoid the ineffective computations resulting from the activation and weight parameters in computationally expensive convolution operations. The bus-specific clock gating is implemented across the registers associated with the adder units in MAC modules to minimize unnecessary signal transitions. This approach utilizes dual-port memory partitioning and memory splitting techniques to enhance memory access efficiency and improve parallelism. The results show that our method achieves a 1.58<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times \)</EquationSource> </InlineEquation> reduction in power consumption, and the design exhibits a significant reduction in resource utilization, with a 2.9<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\times \)</EquationSource> </InlineEquation> reduction in (Look Up Table) LUT utilization, as well as lower usage of BRAM, DSP (Digital Signal Processing) units, and flip-flop resources. The low-power and resource-efficient design makes the proposed work suitable for Edge AI applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual sparsity aware PE networks for CNN accelerators in edge AI deployments

  • Akshayraj M. R.,
  • Muhammed Raees P. C.,
  • Varun P. Gopi,
  • Lakshminarayanan G.,
  • Gangadharan G. R.

摘要

This work proposes a dual-sparsity-aware processing element (PE) incorporated within the systolic array network, which can avoid the ineffective computations resulting from the activation and weight parameters in computationally expensive convolution operations. The bus-specific clock gating is implemented across the registers associated with the adder units in MAC modules to minimize unnecessary signal transitions. This approach utilizes dual-port memory partitioning and memory splitting techniques to enhance memory access efficiency and improve parallelism. The results show that our method achieves a 1.58 \(\times \) reduction in power consumption, and the design exhibits a significant reduction in resource utilization, with a 2.9 \(\times \) reduction in (Look Up Table) LUT utilization, as well as lower usage of BRAM, DSP (Digital Signal Processing) units, and flip-flop resources. The low-power and resource-efficient design makes the proposed work suitable for Edge AI applications.