<p>Convolutional Neural Networks (CNNs) are among the most promising algorithms, outperforming traditional methods in classification tasks with superior accuracy. They have been widely applied across various deep learning domains, including computer vision, speech recognition, image processing, and object detection. However, many CNNs require substantial computational resources, particularly within their convolutional layers. As high-performance CNNs continue to evolve, their processing and memory requirements are also increasing. To address these challenges, this paper proposes an effective design methodology for accelerating CNN algorithms on Field-Programmable Gate Array (FPGA) hardware architectures. The proposed methodology introduces a novel approach for accelerating CNN algorithms using FPGAs, addressing the significant processing and memory demands associated with CNNs. The implementation is based on Open Computing Language (OpenCL), which provides rapid implementation flows. This approach was chosen for its efficiency in reducing development time and eliminating the need to manually write hardware description language (HDL) code. The MNIST and the CIFAR-10 datasets on the Xilinx ZYNQ 7000 device were used to evaluate our approach. Our method achieved a 97% recognition rate on MNIST and an 86% recognition rate on CIFAR-10. We compared the execution time of our accelerated CNN kernel on the FPGA with that of a single-core Central Processing Unit (CPU). The experimental results demonstrate that our proposed design is 10 times faster than a standard CPU, validating its effectiveness. Our model optimizes power consumption and performance, exceeding previous studies in accuracy and efficiency. It is well suited for real-world applications that demand both precision and energy efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accelerating convolutional neural networks on FPGA platforms: a high-performance design methodology using OpenCL

  • Soufien Gdaim,
  • Abdellatif Mtibaa

摘要

Convolutional Neural Networks (CNNs) are among the most promising algorithms, outperforming traditional methods in classification tasks with superior accuracy. They have been widely applied across various deep learning domains, including computer vision, speech recognition, image processing, and object detection. However, many CNNs require substantial computational resources, particularly within their convolutional layers. As high-performance CNNs continue to evolve, their processing and memory requirements are also increasing. To address these challenges, this paper proposes an effective design methodology for accelerating CNN algorithms on Field-Programmable Gate Array (FPGA) hardware architectures. The proposed methodology introduces a novel approach for accelerating CNN algorithms using FPGAs, addressing the significant processing and memory demands associated with CNNs. The implementation is based on Open Computing Language (OpenCL), which provides rapid implementation flows. This approach was chosen for its efficiency in reducing development time and eliminating the need to manually write hardware description language (HDL) code. The MNIST and the CIFAR-10 datasets on the Xilinx ZYNQ 7000 device were used to evaluate our approach. Our method achieved a 97% recognition rate on MNIST and an 86% recognition rate on CIFAR-10. We compared the execution time of our accelerated CNN kernel on the FPGA with that of a single-core Central Processing Unit (CPU). The experimental results demonstrate that our proposed design is 10 times faster than a standard CPU, validating its effectiveness. Our model optimizes power consumption and performance, exceeding previous studies in accuracy and efficiency. It is well suited for real-world applications that demand both precision and energy efficiency.