Swarm Learning (SL) is a decentralized machine learning paradigm that uses blockchain to enable secure collaboration among edge devices without data sharing. However, due to data and resource heterogeneity, SL suffers from the straggler problem, where slow devices delay synchronous aggregation. Existing solutions like load-balanced synchronization are infeasible under SL’s strict privacy constraints, while asynchronous methods struggle with decentralized coordination and dynamic aggregator selection. To address these challenges, we propose Seaswarm, a semi-asynchronous SL framework that balances convergence speed and model accuracy while preserving decentralization. Seaswarm introduces a tiered synchronization mechanism, combining a peer grouping algorithm and non-blocking All-Reduce within groups to alleviate resource imbalance. It also employs a load balancing strategy based on runtime metrics to address data heterogeneity. Experiments on five non-IID datasets across both high-performance and Raspberry Pi clusters show that Seaswarm achieves up to 8.0× speedup over standard SL on servers and 1.23× improvement over FedAT on Raspberry Pi clusters, without sacrificing accuracy. These results demonstrate Seaswarm’s effectiveness in accelerating decentralized training in heterogeneous decentralized learning environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Seaswarm: Mitigating Stragglers in Cross-Silo Swarm Learning via Semi-asynchronous Communication

  • Xiaoyue Yang,
  • Chuang Hu,
  • Dazhao Cheng

摘要

Swarm Learning (SL) is a decentralized machine learning paradigm that uses blockchain to enable secure collaboration among edge devices without data sharing. However, due to data and resource heterogeneity, SL suffers from the straggler problem, where slow devices delay synchronous aggregation. Existing solutions like load-balanced synchronization are infeasible under SL’s strict privacy constraints, while asynchronous methods struggle with decentralized coordination and dynamic aggregator selection. To address these challenges, we propose Seaswarm, a semi-asynchronous SL framework that balances convergence speed and model accuracy while preserving decentralization. Seaswarm introduces a tiered synchronization mechanism, combining a peer grouping algorithm and non-blocking All-Reduce within groups to alleviate resource imbalance. It also employs a load balancing strategy based on runtime metrics to address data heterogeneity. Experiments on five non-IID datasets across both high-performance and Raspberry Pi clusters show that Seaswarm achieves up to 8.0× speedup over standard SL on servers and 1.23× improvement over FedAT on Raspberry Pi clusters, without sacrificing accuracy. These results demonstrate Seaswarm’s effectiveness in accelerating decentralized training in heterogeneous decentralized learning environments.