Advanced HLS for Reconfigurable Image Processing in Big Data Systems
摘要
This paper introduces novel HLS techniques for reconfigurable and memory-efficient image processing within deep learning frameworks, addressing inherent limitations of current deep learning accelerators (DLAs) due to their dependency on the host processor, leading to pipeline delays and reduced throughput. The proposed approach emphasizes parallelization, on-chip buffering, and scalability, enhancing computational efficiency and throughput for AI applications in big data and edge computing environments. Methodologically, the paper emphasizes mathematical analysis, hardware architecture specification, HLS modeling, Verilog RTL implementation. By addressing these challenges, this work contributes to advancing the performance capabilities of DLAs in deep learning inference tasks. The algorithms integrates seamlessly with DLA IP subsystems, such as AXI interfaces, external (DRAM) and internal (SRAM) memory interfaces, and activation/pooling engines. Key contributions include advanced algorithm design to enhance memory efficiency and reconfigurability, essential for adapting to diverse hardware architectures. Experimental validation demonstrates substantial improvements in computational speed and resource utilization compared to traditional methods, making it suitable for intensive AI tasks in big data analytics. In conclusion, this research presents a framework for integrating HLS techniques into hardware accelerators, facilitating efficient image processing across diverse applications in big data and edge computing. These advancements aim to bolster DLAs’ capabilities in handling complex AI workloads effectively, promising enhanced solutions for AI-driven domains.