Designing Resource-Efficient Hardware Arithmetic for FPGA-Based Accelerators Leveraging Approximations and Mixed Quantizations
摘要
While ASIC-based hardware platforms provide better application-specific cost–accuracy trade-offs, the diversity of embedded systems deploying machine learning algorithms has risen steadily. Consequently, given their reconfigurability and high performance, FPGA-based hardware platforms are increasingly used for embedded machine learning. However, the low-power designs devised for ASICs, using methods such as precision scaling, approximate computing, and mixed/custom quantization, do not result in proportionate gains when implemented on FPGAs. This lack of proportional gains can be attributed primarily to the lack of optimizations for FPGA’s LUT-based architecture in the ASIC-optimized designs. Consequently, there has been active research on improving the efficacy of low-power methods in FPGA-based systems. In this chapter, we provide an overview of such FPGA-oriented low-power design methods and delve into the details of selected works that report considerable improvements in this regard. Specifically, we cover custom optimizations for both accurate and approximate multiplier designs and MAC units employing mixed quantization of Posit and fixed-point/integer number representations.