<p><Emphasis Type="Underline">A</Emphasis> <Emphasis Type="Underline">L</Emphasis>ite <Emphasis Type="Underline">B</Emphasis>idirectional <Emphasis Type="Underline">E</Emphasis>ncoder <Emphasis Type="Underline">R</Emphasis>epresentations from <Emphasis Type="Underline">T</Emphasis>ransformers model is demonstrated on an analog inference chip fabricated at 14nm node with phase change memory. The 7.1 million unique analog weights shared across 12 layers are mapped to a single chip, accurately programmed into the conductance of 28.3 million devices, for this first analog hardware demonstration of a meaningfully large Transformer model. The implemented model achieved near iso-accuracy on the General Language Understanding Evaluation benchmark of seven tasks, despite the presence of weight-programming errors, hardware imperfections, readout noise, and error propagation. The average hardware accuracy was only 1.8% below that of the floating-point reference, with several tasks at full iso-accuracy. Careful fine-tuning of model weights using hardware-aware techniques contributes an average hardware accuracy improvement of 4.4%. Accuracy loss due to conductance drift – measured to be roughly 5% over 30 days – was reduced to less than 1% with a recalibration-based “drift compensation” technique.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Demonstration of transformer-based ALBERT model on a 14nm analog AI inference chip

  • An Chen,
  • Stefano Ambrogio,
  • Pritish Narayanan,
  • Atsuya Okazaki,
  • Charles Mackin,
  • Andrea Fasoli,
  • Malte J. Rasch,
  • Alexander Friz,
  • Jose Luquin,
  • Takeo Yasuda,
  • Masatoshi Ishii,
  • Takuto Kanamori,
  • Kohji Hosokawa,
  • Timothy Philicelli,
  • Seiji Munetoh,
  • Vijay Narayanan,
  • Hsinyu Tsai,
  • Geoffrey W. Burr

摘要

A Lite Bidirectional Encoder Representations from Transformers model is demonstrated on an analog inference chip fabricated at 14nm node with phase change memory. The 7.1 million unique analog weights shared across 12 layers are mapped to a single chip, accurately programmed into the conductance of 28.3 million devices, for this first analog hardware demonstration of a meaningfully large Transformer model. The implemented model achieved near iso-accuracy on the General Language Understanding Evaluation benchmark of seven tasks, despite the presence of weight-programming errors, hardware imperfections, readout noise, and error propagation. The average hardware accuracy was only 1.8% below that of the floating-point reference, with several tasks at full iso-accuracy. Careful fine-tuning of model weights using hardware-aware techniques contributes an average hardware accuracy improvement of 4.4%. Accuracy loss due to conductance drift – measured to be roughly 5% over 30 days – was reduced to less than 1% with a recalibration-based “drift compensation” technique.