GLoRia: An Energy-Efficient GPU-RRAM System Stack for Large Neural Networks
摘要
While the potential of in-memory RRAM computation for achieving energy-efficient NNs is recognized, concerns persist about its relative scalability to support modern NNs with billions of parameters. In this context, this paper presents GLoRia, a GPU-RRAM architecture and associated software stack to handle these limitations. We strategically identify the optimal NN layers for RRAM acceleration, enhancing the scalability of RRAMs for complex NN architectures and reducing energy consumption. We validate our approach using practical large CNN and GPT models, showing a 6.4 \(\times \) decrease in energy consumption, without compromising inference accuracy, thanks to the proposed strategy.