Architecture-Compiler Co-design for ReRAM-Based Multi-core CIM Architectures
摘要
Resistive Random-Access Memory (ReRAM)-based multi-core systems improve the performance of Convolutional Neural Network (CNN) inference. However, the potential speedup is limited by two factors, the interconnect of the cores and the workload itself. This paper investigates the impact of these factors on the inference latency of CNNs. The ReRAM-based Computing-in-Memory (CIM) architectures are modeled in a cycle-accurate SystemC Transaction-Level Modeling (TLM)-2.0-based simulator to analyze the impact of different architecture parameters. We further develop a compiler tailored to the architecture to execute and compare different CNN workloads. Depending on the architecture setup and workload, a CIM utilization of up to 95% can be achieved.