Multilayer Multipurpose Caches for OpenMP Target Regions on FPGAs
摘要
Multipurpose caches can improve the throughput between the FPGA’s memory and the hardware that is generated when offloading OpenMP target regions. We discuss and evaluate the weaknesses (and also advantages) of different cacheing techniques in this context. Our OpenMP-to-FPGA compiler fully automatically combines and inserts them as a multilayer cache to get the best of all worlds. We evaluate on a diverse benchmark and achieve an average speedup of 3.65, outperforming 1-layer caches both in terms of runtime and resilience.