<p>The Least Mean Square (LMS) adaptive filtering algorithm is a significant filtering algorithm widely used in noise processing and other fields that automatically adjusts the values of filter coefficients according to the results, aimed at optimizing the filtered results. Based on a basic serial LMS adaptive filtering algorithm, we propose a vectorized parallel processing scheme for the LMS adaptive filtering algorithm in this work. By combining the characteristics of the algorithm processing flow and those of the parallel technologies used in vector Digital Signal Processes (DSPs), the optimizations such as loop fusion, double-word accessing, and vector shuffling of the LMS algorithm are studied in depth, and the loop unrolling optimization method is used to accelerate the calculation of the algorithm further. Experimental research was conducted on the high-performance FT-M7002 DSP platform in this paper. The results show that, compared with the running performance of the LMS adaptive filtering algorithm in Texas Instruments (TI)’s dsplib library on the TMS320C6678 processor, the optimization effect of the proposed optimization algorithm in this paper can achieve a maximum speed-up ratio of up to 6.9<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times \)</EquationSource> </InlineEquation> for medium-scale data. The merged memory access optimization implemented on the GPU platform achieves an average 1.5x speedup compared to the basic parallel scheme.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Parallel implementation and optimization of LMS adaptive filtering algorithms based on vector DSP

  • Yonghua Hu,
  • Linyun Deng,
  • Xiangyu Gao,
  • Zhezhuo Zhao

摘要

The Least Mean Square (LMS) adaptive filtering algorithm is a significant filtering algorithm widely used in noise processing and other fields that automatically adjusts the values of filter coefficients according to the results, aimed at optimizing the filtered results. Based on a basic serial LMS adaptive filtering algorithm, we propose a vectorized parallel processing scheme for the LMS adaptive filtering algorithm in this work. By combining the characteristics of the algorithm processing flow and those of the parallel technologies used in vector Digital Signal Processes (DSPs), the optimizations such as loop fusion, double-word accessing, and vector shuffling of the LMS algorithm are studied in depth, and the loop unrolling optimization method is used to accelerate the calculation of the algorithm further. Experimental research was conducted on the high-performance FT-M7002 DSP platform in this paper. The results show that, compared with the running performance of the LMS adaptive filtering algorithm in Texas Instruments (TI)’s dsplib library on the TMS320C6678 processor, the optimization effect of the proposed optimization algorithm in this paper can achieve a maximum speed-up ratio of up to 6.9 \(\times \) for medium-scale data. The merged memory access optimization implemented on the GPU platform achieves an average 1.5x speedup compared to the basic parallel scheme.