<p>Independent Component Analysis (ICA) is a fundamental algorithm for solving blind source separation (BSS) problems, with FastICA being one of the most widely used and computationally efficient methods. However, its performance in real-time data processing is often limited. The paper proposes an optimized FPGA-based fixed-point streaming architecture to efficiently accelerate FastICA processing. The proposed architecture improves computational performance while minimizing resource consumption. The internal structure of each computational sub-unit has been redesigned and optimized to maximize parallelism. A dataflow transmission model interconnects these subunits, enabling a fully pipelined data processing approach. Furthermore, the architecture integrates stream interfaces, which ensures seamless dataflow operation and significantly reduces the overhead associated with data migration between on-chip and off-chip memory, improving real-time processing capabilities. The proposed accelerator has been successfully implemented on the Xilinx KV260 platform. Performance evaluation using synthetic signals demonstrates that the accelerator can process 4 channel FastICA operations on 1024-sample signals within just 0.407&#xa0;ms at a clock frequency of 100&#xa0;MHz. We also compared the performance of the system for continuous data processing on both FPGA and GPU. Experimental results demonstrated that the design achieved a speedup ranging from 19x to 215x, highlighting its significant performance advantage.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FiSA: Efficient Fixed-Point Stream Architecture for FastICA Implementing on FPGA

  • Lianyou Lai,
  • Yazhe Zhang,
  • Ling Qin,
  • Weijian Xu

摘要

Independent Component Analysis (ICA) is a fundamental algorithm for solving blind source separation (BSS) problems, with FastICA being one of the most widely used and computationally efficient methods. However, its performance in real-time data processing is often limited. The paper proposes an optimized FPGA-based fixed-point streaming architecture to efficiently accelerate FastICA processing. The proposed architecture improves computational performance while minimizing resource consumption. The internal structure of each computational sub-unit has been redesigned and optimized to maximize parallelism. A dataflow transmission model interconnects these subunits, enabling a fully pipelined data processing approach. Furthermore, the architecture integrates stream interfaces, which ensures seamless dataflow operation and significantly reduces the overhead associated with data migration between on-chip and off-chip memory, improving real-time processing capabilities. The proposed accelerator has been successfully implemented on the Xilinx KV260 platform. Performance evaluation using synthetic signals demonstrates that the accelerator can process 4 channel FastICA operations on 1024-sample signals within just 0.407 ms at a clock frequency of 100 MHz. We also compared the performance of the system for continuous data processing on both FPGA and GPU. Experimental results demonstrated that the design achieved a speedup ranging from 19x to 215x, highlighting its significant performance advantage.