DIANA: DIgital and ANAlog Heterogeneous Multi-core System-on-Chip
摘要
This chapter further explores the trade-off between flexibility and energy efficiency by extending the idea of heterogeneity initiated with TinyVers toward multiple diverse accelerators. The resulting DIANA system-on-chip (SoC) is a heterogeneous multi-core accelerator that combines a RISC-V host processor with analog in-memory computing (analog in-memory compute (AIMC)) artificial intelligence (AI) accelerator and a digital reconfigurable deep neural networks (DNN) accelerator in a single system-on-chip to support a wide variety of neural network workloads. AIMC cores can bring extreme computational parallelism and efficiency at the expense of accuracy and dataflow flexibility. Digital AI co-processors, on the other hand, guarantee accuracy through deterministic computing but cannot achieve the same computational density and efficiency. DIANA exploits this fundamental trade-off by integrating both types of cores in a shared and optimized memory system to enable seamless execution of the workloads on the parallel cores. The system’s performance benefits further from pipelined parallel execution across both accelerator cores and enhanced AIMC spatial unrolling techniques, leading to drastically reduced execution latency and reduced memory footprints. The design has been implemented in a 22-nm technology and achieves peak efficiencies of 600 TOP/s/W for the AIMC core (I/W/O: 7/1.5/6bit) and 14 TOP/s/W (I/W/O: 8/8/8bit) for the digital accelerator, respectively. End-to-end performance evaluation of CIFAR-10 and ImageNet classification workloads are carried out on the chip, reporting 7.02 TOP/s/W and 5.56 TOP/s/W, respectively, at the system level.