A Verilog implementation is made towards pattern recognition using convolutional neural network (CNN). The implementation derived demonstrates detection of a given pattern by revealing—‘number of times the particular pattern is being detected’. This is devised using adders, look-up-table (LUT) logics and flip-flops (FF) in a manner which satisfies the CNN architecture. This brief studies about reduction in the resources required for compute intensive processes like processing of the video/ graphics. An implementation of neural network is made using adders in an attempt to perform the compute intensive processes on CPU itself. In this way the requirement of GPUs could be avoided for some of the applications. The Verilog implementation in the devised architecture resulted in overall 11% less resource consumption while running simulations considering a 200 MHz clock frequency. It is observed that devised architecture consumed 2798 LUT count and 1652 FF count as opposed to one of the existing designs which shows a consumption of 10,292 LUT and 10,827 FF.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigation of an Architecture for Convolutional Neural Network with Efficient Resource Utilization

  • Kishan Mishra,
  • Sushanta Bordoloi

摘要

A Verilog implementation is made towards pattern recognition using convolutional neural network (CNN). The implementation derived demonstrates detection of a given pattern by revealing—‘number of times the particular pattern is being detected’. This is devised using adders, look-up-table (LUT) logics and flip-flops (FF) in a manner which satisfies the CNN architecture. This brief studies about reduction in the resources required for compute intensive processes like processing of the video/ graphics. An implementation of neural network is made using adders in an attempt to perform the compute intensive processes on CPU itself. In this way the requirement of GPUs could be avoided for some of the applications. The Verilog implementation in the devised architecture resulted in overall 11% less resource consumption while running simulations considering a 200 MHz clock frequency. It is observed that devised architecture consumed 2798 LUT count and 1652 FF count as opposed to one of the existing designs which shows a consumption of 10,292 LUT and 10,827 FF.