<p>Recent advancements in convolutional neural network (CNN)-based object detection models have significantly improved detection accuracy; however, these improvements have come at the cost of increased computational and memory demands, posing challenges for efficient deployment in resource-constrained edge environments. To address these challenges, existing CNN accelerators have primarily focused on enhancing computational efficiency through the adoption of lightweight models such as MobileNetV1. Nevertheless, systematic analyses of layer-wise memory requirements have been relatively lacking, often resulting in inefficient utilization of on-chip memory (OCM) resources during deployment. To overcome these limitations, this paper proposes <i>Auto-Accel</i>, a SW/HW co-design solution for CNN accelerators that simultaneously enables energy-efficient computation and improves memory resource utilization. On the SW side, <i>Auto-Accel</i> effectively reduces hardware resource utilization by proposing a fused quantization technique based on the sum-of-power-of-two scaling factor approximation. On the HW side, <i>Auto-Accel</i> includes an adaptive unified buffer mapping that efficiently reallocates buffer resources according to the layer-wise memory requirements of activations and weights. Furthermore, to maximize data reuse and reduce off-chip memory accesses, we propose a tile-based adaptive pipelined dataflow, which maximizes computational and energy efficiency. The MobileNetV1-SSD lite accelerator equipped with the proposed <i>Auto-Accel</i> achieves an energy efficiency of 11.5&#xa0;FPS/W when implemented on a ZCU102 board, representing an improvement of approximately 1.33<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> to 4.41<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> in energy efficiency compared to prior studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Auto-accel: a SW/HW co-design framework with adaptive unified buffer mapping and power-of-two quantization for FPGA-based object detection accelerators

  • Junha Ko,
  • Dongjun Lee,
  • Youngchan Kim,
  • Hyun Kim

摘要

Recent advancements in convolutional neural network (CNN)-based object detection models have significantly improved detection accuracy; however, these improvements have come at the cost of increased computational and memory demands, posing challenges for efficient deployment in resource-constrained edge environments. To address these challenges, existing CNN accelerators have primarily focused on enhancing computational efficiency through the adoption of lightweight models such as MobileNetV1. Nevertheless, systematic analyses of layer-wise memory requirements have been relatively lacking, often resulting in inefficient utilization of on-chip memory (OCM) resources during deployment. To overcome these limitations, this paper proposes Auto-Accel, a SW/HW co-design solution for CNN accelerators that simultaneously enables energy-efficient computation and improves memory resource utilization. On the SW side, Auto-Accel effectively reduces hardware resource utilization by proposing a fused quantization technique based on the sum-of-power-of-two scaling factor approximation. On the HW side, Auto-Accel includes an adaptive unified buffer mapping that efficiently reallocates buffer resources according to the layer-wise memory requirements of activations and weights. Furthermore, to maximize data reuse and reduce off-chip memory accesses, we propose a tile-based adaptive pipelined dataflow, which maximizes computational and energy efficiency. The MobileNetV1-SSD lite accelerator equipped with the proposed Auto-Accel achieves an energy efficiency of 11.5 FPS/W when implemented on a ZCU102 board, representing an improvement of approximately 1.33 \(\times \) × to 4.41 \(\times \) × in energy efficiency compared to prior studies.