FederatedHPC: a scalable framework for privacy-preserving federated learning in scientific supercomputing workflows
摘要
High-performance computing (HPC) systems increasingly drive collaborative scientific discovery, where data remains distributed, sensitive, and constrained by institutional policies. Conventional federated learning (FL) methods struggle to meet the demands of such environments due to job preemptions, communication bottlenecks, energy constraints, and strict privacy requirements. This paper presents FederatedHPC, a scalable, privacy-preserving, and resource-aware FL framework tailored for HPC deployments. The framework employs a hierarchical aggregation scheme to reduce cross-cluster communication while preserving statistical accuracy, coupled with elastic checkpointing and staleness-aware updates to handle node volatility. An adaptive synchronization mechanism dynamically balances convergence efficiency and bandwidth utilization. A scalarized multi-objective optimization formulation jointly minimizes model loss, communication cost, energy consumption, and privacy leakage, quantified through a formal resilience metric against inference attacks. Comprehensive simulations on heterogeneous node profiles modeled after Summit, Shaheen II, and SuperMUC-NG supercomputers show that FederatedHPC reduces communication volume by