<p>Because of the performance bottleneck faced by CPU when dealing with massive image data, heterogeneous parallel computing has become one of the research hotspots in the field of parallel computing. In this paper, a method is proposed to improve the processing speed of bilateral filters by using CPU–GPU heterogeneous architecture. Under the condition of OpenCL, the series-parallel analysis of the bilateral filter algorithm is carried out, and the optimization is carried out from the point of view of coarse-grained and fine-grained respectively. The performance of fine-grained spatial weight and grayscale similarity weight is optimized by using the data vectorization processing, and the speed-up effect is compared with the multi-core CPU bilateral filter parallel algorithm and the CUDA bilateral filter parallel algorithm. Local memory and direct rendering technology are used to improve memory bandwidth and CPU–GPU utilization. The results show that when the computing scale reaches tens of millions of magnitude, compared with the CPU serial algorithm, the parallel algorithms on the multi-core CPU platform, and on the CUDA platform, the speedup of the bilateral filter under OpenCL architecture achieved 77.12 times, 13.99 times and 1.2 times, respectively. The computational efficiency is greatly improved and the results are consistent with those of serial computing. It provides a new solution idea for the optimization research of image algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on parallel computing of bilateral filter oriented to OpenCL model

  • Chuanyun Liu,
  • Han Xiao,
  • Xiaoyu Ji,
  • Yupu Song,
  • Qinglei Zhou

摘要

Because of the performance bottleneck faced by CPU when dealing with massive image data, heterogeneous parallel computing has become one of the research hotspots in the field of parallel computing. In this paper, a method is proposed to improve the processing speed of bilateral filters by using CPU–GPU heterogeneous architecture. Under the condition of OpenCL, the series-parallel analysis of the bilateral filter algorithm is carried out, and the optimization is carried out from the point of view of coarse-grained and fine-grained respectively. The performance of fine-grained spatial weight and grayscale similarity weight is optimized by using the data vectorization processing, and the speed-up effect is compared with the multi-core CPU bilateral filter parallel algorithm and the CUDA bilateral filter parallel algorithm. Local memory and direct rendering technology are used to improve memory bandwidth and CPU–GPU utilization. The results show that when the computing scale reaches tens of millions of magnitude, compared with the CPU serial algorithm, the parallel algorithms on the multi-core CPU platform, and on the CUDA platform, the speedup of the bilateral filter under OpenCL architecture achieved 77.12 times, 13.99 times and 1.2 times, respectively. The computational efficiency is greatly improved and the results are consistent with those of serial computing. It provides a new solution idea for the optimization research of image algorithms.