<p>Artificial intelligence (AI) edge devices<sup><CitationRef AdditionalCitationIDS="CR2 CR3 CR4 CR5 CR6 CR7 CR8 CR9 CR10 CR11" CitationID="CR1">1</CitationRef>–<CitationRef CitationID="CR12">12</CitationRef></sup> demand high-precision energy-efficient computations, large on-chip model storage, rapid wakeup-to-response time and cost-effective foundry-ready solutions. Floating&#xa0;point (FP) computation provides precision exceeding that of integer (INT) formats at the cost of higher power and storage overhead. Multi-level-cell (MLC) memristor compute-in-memory (CIM)<sup><CitationRef AdditionalCitationIDS="CR14" CitationID="CR13">13</CitationRef>–<CitationRef CitationID="CR15">15</CitationRef></sup> provides compact non-volatile storage and energy-efficient computation but is prone to accuracy loss owing to process variation. Digital static random-access memory (SRAM)-CIM<sup><CitationRef AdditionalCitationIDS="CR17 CR18 CR19 CR20 CR21" CitationID="CR16">16</CitationRef>–<CitationRef CitationID="CR22">22</CitationRef></sup> enables lossless computation; however, storage is low as a result of large bit-cell area and model loading is required during inference. Thus, conventional approaches using homogeneous CIM architectures and computation formats impose a trade-off between efficiency, storage, wakeup latency and inference accuracy. Here we present a mixed-precision heterogeneous CIM AI edge processor, which supports the layer-granular/kernel-granular partitioning of network layers among on-chip CIM architectures (that is, memristor-CIM, SRAM-CIM and tiny-digital units) and computation number formats (INT and FP) based on sensitivity to error. This layer-granular/kernel-granular flexibility allows simultaneous optimization within the two-dimensional design space at the hardware level. The proposed hardware achieved high energy efficiency (40.91 TFLOPS W<sup>−1</sup> for ResNet-20 with CIFAR-100 and 28.63 TFLOPS W<sup>−1</sup> for MobileNet-v2 with ImageNet), low accuracy degradation (&lt;0.45% for ResNet-20 with CIFAR-100 and for MobilNet-v2 with ImageNet) and rapid&#xa0;wakeup-to-response time (373.52 μs).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A mixed-precision memristor and SRAM compute-in-memory AI processor

  • Win-San Khwa,
  • Tai-Hao Wen,
  • Hung-Hsi Hsu,
  • Wei-Hsing Huang,
  • Yu-Chen Chang,
  • Ting-Chien Chiu,
  • Zhao-En Ke,
  • Yu-Hsiang Chin,
  • Hua-Jin Wen,
  • Wei-Ting Hsu,
  • Chung-Chuan Lo,
  • Ren-Shuo Liu,
  • Chih-Cheng Hsieh,
  • Kea-Tiong Tang,
  • Mon-Shu Ho,
  • Ashwin Sanjay Lele,
  • Shih-Hsin Teng,
  • Chung-Cheng Chou,
  • Yu-Der Chih,
  • Tsung-Yung Jonathan Chang,
  • Meng-Fan Chang

摘要

Artificial intelligence (AI) edge devices112 demand high-precision energy-efficient computations, large on-chip model storage, rapid wakeup-to-response time and cost-effective foundry-ready solutions. Floating point (FP) computation provides precision exceeding that of integer (INT) formats at the cost of higher power and storage overhead. Multi-level-cell (MLC) memristor compute-in-memory (CIM)1315 provides compact non-volatile storage and energy-efficient computation but is prone to accuracy loss owing to process variation. Digital static random-access memory (SRAM)-CIM1622 enables lossless computation; however, storage is low as a result of large bit-cell area and model loading is required during inference. Thus, conventional approaches using homogeneous CIM architectures and computation formats impose a trade-off between efficiency, storage, wakeup latency and inference accuracy. Here we present a mixed-precision heterogeneous CIM AI edge processor, which supports the layer-granular/kernel-granular partitioning of network layers among on-chip CIM architectures (that is, memristor-CIM, SRAM-CIM and tiny-digital units) and computation number formats (INT and FP) based on sensitivity to error. This layer-granular/kernel-granular flexibility allows simultaneous optimization within the two-dimensional design space at the hardware level. The proposed hardware achieved high energy efficiency (40.91 TFLOPS W−1 for ResNet-20 with CIFAR-100 and 28.63 TFLOPS W−1 for MobileNet-v2 with ImageNet), low accuracy degradation (<0.45% for ResNet-20 with CIFAR-100 and for MobilNet-v2 with ImageNet) and rapid wakeup-to-response time (373.52 μs).