<p>High-performance computing (HPC) systems increasingly drive collaborative scientific discovery, where data remains distributed, sensitive, and constrained by institutional policies. Conventional federated learning (FL) methods struggle to meet the demands of such environments due to job preemptions, communication bottlenecks, energy constraints, and strict privacy requirements. This paper presents FederatedHPC, a scalable, privacy-preserving, and resource-aware FL framework tailored for HPC deployments. The framework employs a hierarchical aggregation scheme to reduce cross-cluster communication while preserving statistical accuracy, coupled with elastic checkpointing and staleness-aware updates to handle node volatility. An adaptive synchronization mechanism dynamically balances convergence efficiency and bandwidth utilization. A scalarized multi-objective optimization formulation jointly minimizes model loss, communication cost, energy consumption, and privacy leakage, quantified through a formal resilience metric against inference attacks. Comprehensive simulations on heterogeneous node profiles modeled after Summit, Shaheen II, and SuperMUC-NG supercomputers show that FederatedHPC reduces communication volume by <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\approx\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≈</mo> </math></EquationSource> </InlineEquation> 35.9 % and energy consumption by <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\approx\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≈</mo> </math></EquationSource> </InlineEquation> 25% compared to baseline methods, without compromising model utility or privacy. The framework further incorporates a fairness-aware regularizer and incentive-driven coordination mechanism to promote equitable and sustainable multi-institution participation. Finally, continual and dynamic training extensions, together with scalability analysis up to 2,048 nodes, demonstrate the framework’s adaptability to evolving workloads and large-scale HPC environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FederatedHPC: a scalable framework for privacy-preserving federated learning in scientific supercomputing workflows

  • Waseem Abbass,
  • Nasim Abbas,
  • Uzma Majeed,
  • Muhammad Ehtisham Hassan

摘要

High-performance computing (HPC) systems increasingly drive collaborative scientific discovery, where data remains distributed, sensitive, and constrained by institutional policies. Conventional federated learning (FL) methods struggle to meet the demands of such environments due to job preemptions, communication bottlenecks, energy constraints, and strict privacy requirements. This paper presents FederatedHPC, a scalable, privacy-preserving, and resource-aware FL framework tailored for HPC deployments. The framework employs a hierarchical aggregation scheme to reduce cross-cluster communication while preserving statistical accuracy, coupled with elastic checkpointing and staleness-aware updates to handle node volatility. An adaptive synchronization mechanism dynamically balances convergence efficiency and bandwidth utilization. A scalarized multi-objective optimization formulation jointly minimizes model loss, communication cost, energy consumption, and privacy leakage, quantified through a formal resilience metric against inference attacks. Comprehensive simulations on heterogeneous node profiles modeled after Summit, Shaheen II, and SuperMUC-NG supercomputers show that FederatedHPC reduces communication volume by \(\approx\) 35.9 % and energy consumption by \(\approx\) 25% compared to baseline methods, without compromising model utility or privacy. The framework further incorporates a fairness-aware regularizer and incentive-driven coordination mechanism to promote equitable and sustainable multi-institution participation. Finally, continual and dynamic training extensions, together with scalability analysis up to 2,048 nodes, demonstrate the framework’s adaptability to evolving workloads and large-scale HPC environments.