<p>Advances in neural network (NN) models and learning methods have resulted in breakthroughs in various fields. A larger NN model is more difficult to install on a computer with limited computing resources. One method for compressing NN models is to quantize the weights, in which the connection weights of the NNs are approximated with low-bit precision. The existing quantization methods for NN models can be categorized into two approaches: quantization-aware training (QAT) and post-training quantization (PTQ). In this study, we focused on the performance degradation of NN models using PTQ. This paper proposes a method for visually evaluating the performance of quantized NNs using topological data analysis (TDA). Subjecting the structure of NNs to TDA allows the performance of quantized NNs to be assessed without experiments or simulations. We developed a TDA-based evaluation method for NNs with low-bit weights by referring to previous research on a TDA-based evaluation method for NNs with high-bit weights. We also tested the TDA-based method using the MNIST dataset. Finally, we compared the performance of the quantized NNs generated by static and dynamic quantization through a visual demonstration.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A TDA-based performance analysis for neural networks with low-bit weights

  • Yugo Ogio,
  • Naoki Tsubone,
  • Yuki Minami,
  • Masato Ishikawa

摘要

Advances in neural network (NN) models and learning methods have resulted in breakthroughs in various fields. A larger NN model is more difficult to install on a computer with limited computing resources. One method for compressing NN models is to quantize the weights, in which the connection weights of the NNs are approximated with low-bit precision. The existing quantization methods for NN models can be categorized into two approaches: quantization-aware training (QAT) and post-training quantization (PTQ). In this study, we focused on the performance degradation of NN models using PTQ. This paper proposes a method for visually evaluating the performance of quantized NNs using topological data analysis (TDA). Subjecting the structure of NNs to TDA allows the performance of quantized NNs to be assessed without experiments or simulations. We developed a TDA-based evaluation method for NNs with low-bit weights by referring to previous research on a TDA-based evaluation method for NNs with high-bit weights. We also tested the TDA-based method using the MNIST dataset. Finally, we compared the performance of the quantized NNs generated by static and dynamic quantization through a visual demonstration.