Echo state networks, as a popular form of reservoir computing models, are recurrent neural networks that consist of three layers, and only the output layer needs to be trained. Compared to other recurrent neural network models, echo state networks offer comparable performance for many tasks but lead to reduced computational requirements, making them highly suitable for resource-constrained edge implementations. In this paper, we present and compare two streamlining dataflow accelerators for echo state networks on FPGA. Our accelerators rely on a direct logic implementation style with fully unrolled computations and are fully quantized to avoid floating-point calculations. After proposing a taxonomy of DNN implementation styles on FPGA, we present a tool flow to automatically create such FPGA accelerators for echo state networks. We then discuss two accelerator versions, one employing DSP blocks available in FPGAs and the other one converting multiplications into add/shift operations mapped solely to LUTs. We evaluate our accelerators on three time-series forecasting tasks and report the achieved accuracy as well as the resource requirements, the latency, the throughput, and the required energy. Our designs achieve extreme-throughput and ultra-low latency, reaching up to 100 Megasamples/s and 9.5 ns, respectively, which significantly outperforms related FPGA-based implementations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ultra-Low Latency and Extreme-Throughput Echo State Neural Networks on FPGA

  • Atousa Jafari,
  • Marco Platzner

摘要

Echo state networks, as a popular form of reservoir computing models, are recurrent neural networks that consist of three layers, and only the output layer needs to be trained. Compared to other recurrent neural network models, echo state networks offer comparable performance for many tasks but lead to reduced computational requirements, making them highly suitable for resource-constrained edge implementations. In this paper, we present and compare two streamlining dataflow accelerators for echo state networks on FPGA. Our accelerators rely on a direct logic implementation style with fully unrolled computations and are fully quantized to avoid floating-point calculations. After proposing a taxonomy of DNN implementation styles on FPGA, we present a tool flow to automatically create such FPGA accelerators for echo state networks. We then discuss two accelerator versions, one employing DSP blocks available in FPGAs and the other one converting multiplications into add/shift operations mapped solely to LUTs. We evaluate our accelerators on three time-series forecasting tasks and report the achieved accuracy as well as the resource requirements, the latency, the throughput, and the required energy. Our designs achieve extreme-throughput and ultra-low latency, reaching up to 100 Megasamples/s and 9.5 ns, respectively, which significantly outperforms related FPGA-based implementations.