<p>Recent advancements in artificial intelligence have significantly propelled edge intelligence applications, particularly in smart factories. Deploying computation-intensive convolutional neural networks on resource-constrained edge devices traditionally relies on offloading to remote clouds or optimizing local computation. However, cloud-assisted methods face issues like unreliable networks and high latency, while local computation is limited by device capabilities. To address these issues, we propose a data-parallel inference method based on local clusters which can maximize the utilization of computational resources on each device. Then, the quantization compression is introduced to edge-padding data for reducing communication overhead. Furthermore, to solve the problem of declining inference accuracy caused by new data, we establish a distributed incremental training and inference framework. This framework dynamically monitors the performance and conducts timely training on new data to ensure sustained improvement in inference accuracy. The experimental results prove that our proposed inference strategy outperforms local inference in terms of performance across various device scales. Additionally, the results validate the effectiveness of the collaborative framework integrating distributed training and inference. Compared with other methods, our inference strategy consistently achieves the lowest latency across 2, 3, and 4 follower node configurations, delivering up to 9.42% reduction in inference time. Compared to the local inference time, our approach can reduce the total inference time by 33.60%, 53.12% and 53.89% on 2, 3 and 4 follower nodes, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient distributed training and inference based on data and feature map parallelism via leveraging edge clusters

  • Hui Guo,
  • Chenfeng Zhang,
  • Xiaogang Wang,
  • Jiaqi Fei,
  • Jiayi Li,
  • Yuze Du,
  • Jian Cao

摘要

Recent advancements in artificial intelligence have significantly propelled edge intelligence applications, particularly in smart factories. Deploying computation-intensive convolutional neural networks on resource-constrained edge devices traditionally relies on offloading to remote clouds or optimizing local computation. However, cloud-assisted methods face issues like unreliable networks and high latency, while local computation is limited by device capabilities. To address these issues, we propose a data-parallel inference method based on local clusters which can maximize the utilization of computational resources on each device. Then, the quantization compression is introduced to edge-padding data for reducing communication overhead. Furthermore, to solve the problem of declining inference accuracy caused by new data, we establish a distributed incremental training and inference framework. This framework dynamically monitors the performance and conducts timely training on new data to ensure sustained improvement in inference accuracy. The experimental results prove that our proposed inference strategy outperforms local inference in terms of performance across various device scales. Additionally, the results validate the effectiveness of the collaborative framework integrating distributed training and inference. Compared with other methods, our inference strategy consistently achieves the lowest latency across 2, 3, and 4 follower node configurations, delivering up to 9.42% reduction in inference time. Compared to the local inference time, our approach can reduce the total inference time by 33.60%, 53.12% and 53.89% on 2, 3 and 4 follower nodes, respectively.