HT-NoC: Reconfigurable High Throughput Network-on-Chip for AI Dataflow Accelerators
摘要
Fully Connected (FC) layers are a bottleneck for many Deep Neural Networks (DNN) algorithms due to their high bandwidth requirements, which makes their hardware acceleration particularly challenging. In this paper, we address this challenge from a communication-centric approach. We propose HT-NoC (High Throughput Network-on-Chip), a reconfigurable NoC to accelerate FC layers. HT-NoC features reconfigurable router connections that adapt to varying data traffic patterns. This enables the utilization of available unused bandwidth and router resources to transport more packets simultaneously. Compared to a baseline mesh NoC, HT-NoC achieves a \(4\times \) reduction in latency and a \(2.7\times \) decrease in energy consumption in the propagation of time-consuming FC layer weight parameters. HT-NoC also achieves favorable performance for some Convolution (CONV) layers. When integrated into an AI dataflow accelerator, HT-NoC achieves a \(3\times \) speedup in executing Feed Forward Network (FFN) blocks in Transformers, outperforming state-of-the-art (SoA) systolic array (SA) based accelerators.