One-to-Many Communication Primitives in Dragonfly Networks
摘要
Collective communication primitives (CCPs), such as multicast and broadcast, are essential for many parallel and distributed applications. In response, this study compares a topology-oblivious algorithm underlying the implementation of CCPs in standard instances of MPI with two topology-aware implementations, based on the LLF and GLF algorithms, and an ideal hardware-assisted approach. Our study reveals workload-dependent performance variations among CCP implementations and highlights the importance of CCP algorithm selection in optimizing application performance in supercomputing environments.