A dual-domain compute-in-memory system for general neural network inference
摘要
Analogue compute-in-memory systems can offer superior energy efficiency and parallelism than conventional digital systems. However, complex regression tasks that require precise floating-point (FP) computing remain challenging with such hardware, and previous approaches have, thus, typically focused on classification tasks requiring low data precision and a limited dynamic range. Here we describe an analogue–digital unified compute-in-memory architecture for general neural network inference. The approach is based on a low-cost dual-domain FP processor and merges analogue compute-in-memory arrays with digital cores. It exhibits a 39.2 times higher energy efficiency than common FP-32 multipliers during FP neural network inference. We use this architecture to develop a memristor-based computing system and illustrate its capabilities with a fully hardware-implemented complex regression task using YOLO. The system exhibits a 2.7 times higher mean average precision (increasing from 0.27 to 0.724, mAP-50) compared with pure analogue compute-in-memory systems.