Slide FFT on a homogeneous mesh in wafer-scale computing
摘要
Searches for signals at low signal-to-noise ratios frequently involve correlations evaluated by the Fast Fourier Transform (FFT). To accelerate the discovery power of present and next-generation multi-messenger observatories, we here explore the implementation of FFT on wafer-scale engines. To minimize the memory overhead of the inherently non-local FFT algorithm on a homogeneous mesh of Processing Elements (PEs) with no global memory on the chip, we introduce a new synchronous slide operation (Slide) exploiting fast interconnect between adjacent PEs. The feasibility of compute-limited performance is demonstrated in linear scaling of Slide execution times with varying array sizes in preliminary benchmarks on the CS-2 WSE. As a first step, this benchmark appears promising for the proposed implementation of high-throughput FFT-based signal processing in multi-messenger astronomy.