<p>Analogue compute-in-memory systems can offer superior energy efficiency and parallelism than conventional digital systems. However, complex regression tasks that require precise floating-point (FP) computing remain challenging with such hardware, and previous approaches have, thus, typically focused on classification tasks requiring low data precision and a limited dynamic range. Here we describe an analogue–digital unified compute-in-memory architecture for general neural network inference. The approach is based on a low-cost dual-domain FP processor and merges analogue compute-in-memory arrays with digital cores. It exhibits a 39.2 times higher energy efficiency than common FP-32 multipliers during FP neural network inference. We use this architecture to develop a memristor-based computing system and illustrate its capabilities with a fully hardware-implemented complex regression task using YOLO. The system exhibits a 2.7 times higher mean average precision (increasing from 0.27 to 0.724, mAP-50) compared with pure analogue compute-in-memory systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A dual-domain compute-in-memory system for general neural network inference

  • Ze Wang,
  • Ruihua Yu,
  • Zhiping Jia,
  • Zhifan He,
  • Tianhao Yang,
  • Bin Gao,
  • Yang Li,
  • Zhenping Hu,
  • Zhenqi Hao,
  • Yunrui Liu,
  • Jianghai Lu,
  • Peng Yao,
  • Jianshi Tang,
  • Qi Liu,
  • He Qian,
  • Huaqiang Wu

摘要

Analogue compute-in-memory systems can offer superior energy efficiency and parallelism than conventional digital systems. However, complex regression tasks that require precise floating-point (FP) computing remain challenging with such hardware, and previous approaches have, thus, typically focused on classification tasks requiring low data precision and a limited dynamic range. Here we describe an analogue–digital unified compute-in-memory architecture for general neural network inference. The approach is based on a low-cost dual-domain FP processor and merges analogue compute-in-memory arrays with digital cores. It exhibits a 39.2 times higher energy efficiency than common FP-32 multipliers during FP neural network inference. We use this architecture to develop a memristor-based computing system and illustrate its capabilities with a fully hardware-implemented complex regression task using YOLO. The system exhibits a 2.7 times higher mean average precision (increasing from 0.27 to 0.724, mAP-50) compared with pure analogue compute-in-memory systems.