<p>Incremental learning for resource-constrained systems has been noticed for its ability to incorporate new learning. This paper presents accelerator for incremental learning with intelligent selection, a low-latency hardware implementation for incremental learning based on a hardware-modified model. The model consists of two components: (1) feature extractor using convolutional layers (2) classifier based on a customized K-nearest neighbor (KNN) algorithm. We modified classifier to reduce resource usage by utilizing two groups of data: (1) data closest to the class mean for general features (2) boundary data for distinguishing features. The implementation focuses on two strategies: first, simultaneous feature computation for minimal latency, and second, processing features in batches to reduce resource consumption. The first strategy reduced latency by 5 to 912 times compared to previous works. In the second strategy, LUT and DSP usage dropped by 4.8 and 21 times, respectively, and decreased delay was also observed in most cases. Despite these improvements, accuracy remained nearly the same as prior methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AILIS: effective hardware accelerator for incremental learning with intelligent selection in classification

  • Nafiseh HosseinpourFardi,
  • Bijan Alizadeh

摘要

Incremental learning for resource-constrained systems has been noticed for its ability to incorporate new learning. This paper presents accelerator for incremental learning with intelligent selection, a low-latency hardware implementation for incremental learning based on a hardware-modified model. The model consists of two components: (1) feature extractor using convolutional layers (2) classifier based on a customized K-nearest neighbor (KNN) algorithm. We modified classifier to reduce resource usage by utilizing two groups of data: (1) data closest to the class mean for general features (2) boundary data for distinguishing features. The implementation focuses on two strategies: first, simultaneous feature computation for minimal latency, and second, processing features in batches to reduce resource consumption. The first strategy reduced latency by 5 to 912 times compared to previous works. In the second strategy, LUT and DSP usage dropped by 4.8 and 21 times, respectively, and decreased delay was also observed in most cases. Despite these improvements, accuracy remained nearly the same as prior methods.